Skip to content
Evidilya, People & Research First®
Submit an RFP
Insight · Dec 2025 · 3 min

Verifiability over explainability: how to govern AI in evidence generation

The useful question for clinical AI is not how the model learned, but whether its outputs can be verified, stress-tested and reproduced. A governance stance that matches what the EU AI Act and FDA now ask for.

People & Research First®

Insight · · 3 min

Most of the debate about AI in clinical research fixates on explainability: opening the model, tracing the weights, narrating how an output was reached. It is an understandable instinct and, for the systems now entering evidence generation, largely the wrong emphasis. We do not ask a radiologist to narrate the neural pathways behind a reading. We verify the reading against a reference standard, test the reader against varied and adversarial cases, and record the conditions under which their judgement holds. The same logic should govern AI in evidence generation, and the regulatory texts published since 2024 have moved decisively in that direction.

Two of those texts set the frame. The EU Artificial Intelligence Act, Regulation (EU) 2024/1689, entered into force on 1 August 2024, with obligations phasing in over the following years and the bulk of the high-risk regime applying from August 2026. Read it closely and what it demands of a high-risk system is not a mechanistic account of learning: it is risk management, data governance, technical documentation, logging, human oversight, and demonstrated accuracy and robustness. Every one of those is a verification obligation. In January 2025 the FDA published draft guidance on the use of artificial intelligence to support regulatory decision-making for drug and biological products, built around a risk-based credibility assessment framework anchored to a defined context of use. Again: not how the model works, but whether it is credible for this specific decision.

Context of use is the load-bearing concept, and it is the one sponsors most often skip. A model is never simply validated. It is validated for a stated purpose, on a stated population, producing a stated output that carries a stated weight in a stated decision. The same model used to prioritise a monitoring queue and to adjudicate an endpoint carries entirely different risk, and therefore entirely different evidentiary requirements. A credibility argument that does not name the decision the output feeds is not a credibility argument.

From that, four practical requirements follow, and they map onto obligations that already exist in this industry rather than inventing new ones.

First, provenance of the training and evaluation data. Which sources, which period, which population, which exclusions, and how the evaluation set was held separate. Without this, no claim about generalisation can be assessed, and drift cannot be detected later because there is no baseline to compare against.

Second, adversarial and subgroup stress-testing. Aggregate accuracy conceals the failures that matter. Performance has to be reported across the subgroups the study actually contains, including the small ones, and against deliberately difficult cases rather than a convenience sample. A model that performs well on the median participant and poorly on the participants a study was designed to reach has failed at exactly the point of interest.

Third, human oversight that is real rather than nominal. ICH E6(R3), adopted at Step 4 in January 2025, is deliberately technology-neutral and puts risk-proportionate quality management at the centre of conduct. Applied to AI, that means a named human accountable for the output, with the authority and the information to overrule it, and a record of when they did. Oversight that cannot be exercised is not oversight; it is a signature.

Fourth, an audit trail that reconstructs the decision. Which model version, which inputs, which output, which human reviewed it and when. This is not a novel demand: 21 CFR Part 11 and the computerised-systems expectations in E6(R3) already require it of any system touching regulated data. AI does not get an exemption because its internals are harder to describe.

None of this rules out explainability work, and there are places where it earns its keep, particularly in model development and in debugging failures. The objection is to treating it as the governance answer. A convincing post-hoc explanation of a wrong output is worse than no explanation, because it manufactures confidence. Verification does the opposite: it establishes the conditions under which the output can be trusted, and marks the boundary past which it cannot.

This is the stance behind Evidilya's AI governance. Every AI-assisted step in evidence generation is scoped to a declared context of use, logged, human-reviewed, and testable against pre-specified expectations. It is a more honest account of what current methods can demonstrate, and a more useful one to a reviewer, who is not assessing the architecture but deciding whether the output holds.

Sources. Regulation (EU) 2024/1689 (Artificial Intelligence Act), in force 1 August 2024. FDA, Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products, draft guidance, January 2025. ICH E6(R3) Good Clinical Practice, Step 4, 6 January 2025. 21 CFR Part 11.

Byline

Written by the Evidilya scientific team. For interviews, references or a full publication list, use the contact page.

Talk to our team.

Bring us the evidence question you're wrestling with, we'll tell you the shortest defensible path.