Skip to content
Evidilya, People & Research First®
Submit an RFP
Insight · Jul 2026 · 4 min

AI in clinical research: from acceleration to methodological rigour

Synthetic control arms, cohort selection, risk-based monitoring: AI is already inside modern trials. Where EMA, FDA and the EU AI Act are drawing the line, and what CROs are expected to validate before letting a model near a protocol.

People & Research First®

Insight · · 4 min

Artificial intelligence entered clinical research through the back door, as an efficiency argument. Screen more records, find the cohort faster, flag the sites that need a visit. That framing was accurate for a while and is now insufficient, because the same techniques have moved from the periphery of a trial into positions where they influence what the evidence says. Once a model helps decide who is included, what a comparator looks like, or which safety signal is escalated, it stops being a productivity tool and becomes part of the method.

Three uses are already routine enough to be worth naming precisely.

Cohort selection and feasibility. Models trained on electronic health records and registry data identify eligible populations, estimate accrual and expose inclusion criteria that quietly exclude the intended population. The gain is real, and so is the risk: a model trained on the patients a health system happens to see will reproduce that system's referral patterns, and a feasibility estimate that inherits them will be confidently wrong in exactly the settings where recruitment is hardest.

Synthetic and external control arms. Constructing a comparator from historical trial data, registry cohorts or routinely collected records is attractive where randomisation is difficult to justify or impossible to operate. It also carries emerging regulatory precedent rather than settled acceptance: regulators have accepted external comparators in specific contexts, typically single-arm oncology and rare disease with well-characterised natural history, and have rejected them where outcome definitions, standard of care or measurement practice differ between the trial and the source. Treating this as an established route is the most common overreach in the field. Treating it as unavailable is the second.

Risk-based monitoring and quality oversight. ICH E6(R3) puts proportionate, risk-based quality management at the centre of trial conduct, which is precisely the space where anomaly detection and centralised statistical monitoring earn their keep. This is the least contested application, because the model prioritises human attention rather than replacing a judgement.

The regulatory picture is no longer sparse. The EMA published a reflection paper on the use of artificial intelligence across the medicinal product lifecycle, setting a risk-based expectation that scales with the impact a model has on patients or on the reliability of the evidence, and encouraging early qualification advice for high-impact uses. The FDA has issued draft guidance on AI used to support regulatory decision-making for drugs and biologics, built around a context-of-use framework: the same model carries different evidentiary expectations depending on what it is being relied on to do. And Regulation (EU) 2024/1689, the EU AI Act, entered into force in August 2024 with obligations phasing in through 2026 and 2027, adding horizontal requirements on risk management, data governance, technical documentation, logging and human oversight that apply irrespective of the medicines framework.

What follows from all three is convergent, and it is not a demand for novel science. It is a demand for the discipline the sector already applies to instruments and assays, applied to models.

Context of use, stated before the model is built. What decision does the output inform, who acts on it, and what happens if it is wrong? Everything else follows from this sentence, and a programme that cannot write it is not ready to validate anything.

Data provenance and representativeness. Where the training data came from, which populations it covers, how missingness behaves, and where the deployment population differs from it. Performance reported without this is a number without a denominator.

Pre-specification. Model configuration, thresholds and the analysis plan fixed before the data is seen, with any adaptation described in advance. A model tuned after looking at outcomes is a hypothesis, not a result.

Human oversight that is real. A named accountable reviewer, an escalation path, and a record of overrides. Oversight that cannot be evidenced did not happen.

Monitoring after deployment. Drift is not an exception, it is the expected behaviour of a model meeting a population that keeps changing. Detection thresholds and a retraining decision rule belong in the plan, not in the incident report.

An audit trail a reviewer can follow. Versioned model, versioned data, logged inputs and outputs, and the ability to reconstruct why a given case was classified as it was. This is the requirement that most often exposes a pilot as unshippable.

The honest position is that AI raises the methodological bar rather than lowering it. A model can find a cohort in days that manual feasibility would take months to characterise, and it can also encode a selection effect so smoothly that nobody notices until the analysis will not replicate. The efficiency is real; it is only worth having if the evidence survives review.

For a CRO the obligation is straightforward to state and demanding to meet: know which models touch the evidence, document what each one is relied on to do, validate against that context, and be able to show the whole chain to an inspector who was not in the room. Acceleration is the easy half. Rigour is what makes it usable.

Sources. EMA reflection paper on the use of artificial intelligence in the medicinal product lifecycle. FDA draft guidance on considerations for the use of artificial intelligence to support regulatory decision-making for drug and biological products. Regulation (EU) 2024/1689 (AI Act), in force August 2024, obligations phasing in 2026 – 2027. ICH E6(R3) Good Clinical Practice.

Byline

Written by the Evidilya scientific team. For interviews, references or a full publication list, use the contact page.

Talk to our team.

Bring us the evidence question you're wrestling with, we'll tell you the shortest defensible path.