Regulated AI
What Evidence Should a GxP AI Supplier Provide?
A supplier cannot validate a regulated customer's use of its own product, because validation is a statement about fitness for a specific intended use, and the supplier does not own that intended use. What a supplier can and should provide is the raw material the customer needs to make that determination itself: documented facts about what the model is, how it was built, how it behaves, what changes without notice, and what happens if the relationship ends. Where a supplier's response to due diligence is reassurance rather than evidence, the gap does not disappear; it transfers silently to the customer's own validation package, where it surfaces later as an unsupported claim.
Model and training-data documentation
The customer needs enough information about the model itself to judge whether its design is plausible for the intended use, independent of the supplier's own performance claims:
- the general model type and architecture class, to the extent the supplier is willing to disclose it, sufficient to understand what kind of failure modes are plausible;
- the nature and provenance of training data: its general source, domain, and any known limitations in representativeness for the customer's population, geography or use case;
- known constraints on the model's applicable scope, including populations, input formats or conditions under which performance has not been established;
- version identity for the specific model instance in production, distinct from the general product name, since "the same product" can mean a materially different model over time.
The EMA's reflection paper on the use of artificial intelligence across the medicinal product lifecycle treats data quality, representativeness and transparency about model limitations as central to whether AI-supported outputs can be relied upon: the customer cannot assess representativeness for its own population without the supplier disclosing what population the model was built and evaluated against in the first place.
Performance evidence under relevant conditions
Generic accuracy figures from a supplier's own benchmark are a starting point, not sufficient evidence. Material performance evidence should be evaluated, or at least evaluable, under conditions that resemble the customer's actual use:
- performance metrics relevant to the specific task the customer will rely on, not only aggregate accuracy across all use cases the product supports;
- known failure modes and the conditions under which the model is more likely to be wrong, incomplete or overconfident;
- evidence of evaluation against data resembling the customer's own domain, or an honest statement that this has not been done, rather than silence on the point;
- for any bias-sensitive use case, evidence that the supplier has evaluated performance across relevant subpopulations rather than only in aggregate.
Where a supplier cannot or will not provide task-relevant performance evidence, the customer's own validation activity has to close that gap directly, typically through independent testing on representative data before relying on the output for any GxP decision. This should be a deliberate, resourced activity in the validation plan, not an assumption inherited from the supplier's marketing material.
Change-notification and lifecycle commitments
An AI capability's behaviour is not fixed at go-live in the way a deterministic system's configuration is. A supplier should commit, contractually, to notifying the customer of:
- retraining or model version changes that could alter output behaviour, distinguished from routine infrastructure or interface changes that do not;
- changes to the underlying training data, prompt structure, or any component the customer's validation evidence depended on;
- known degradation, incidents or discovered failure modes affecting production behaviour, on a timescale that allows the customer to assess impact before it becomes an inspection finding rather than after;
- deprecation or end-of-life timelines for the specific model version in use, with adequate notice to requalify or transition.
The joint FDA-EMA guiding principles for good AI practice, published in January 2026, frame this kind of lifecycle transparency and human oversight as foundational to safe AI use across the product lifecycle rather than a one-time gate at deployment: the expectation is continuous, not a single qualification event. A supplier agreement silent on change notification effectively asks the customer to either retest continuously at its own expense with no advance warning, or to accept undetected model drift as a cost of using the product.
Security, logging and evidence generation
For the customer's own audit-trail and investigation obligations to be met, the supplier's platform needs to generate and expose evidence, not merely process requests:
- logging sufficient to reconstruct which model version produced a given output, with what input, at what time, and under whose authenticated access;
- security controls appropriate to the sensitivity of the data processed, including any subprocessors or downstream services the supplier itself relies on;
- a clear statement of what evidence the customer can access directly versus what exists only in the supplier's own systems and would require a request, along with the expected response time for that request, since an inspection timeline will not wait indefinitely.
A model that cannot identify which version produced a historical output leaves the customer unable to investigate a later-discovered defect against the specific behaviour that actually occurred, which is a data-integrity gap regardless of how well the model performs on average.
What must remain the customer's responsibility
No amount of supplier evidence transfers the underlying accountability. The customer, not the supplier, is responsible for:
- defining the intended use the AI capability will be relied upon for, and the boundary of permitted reliance;
- classifying the use case's risk and determining the proportionate validation, monitoring and human-oversight controls;
- deciding whether the supplier's evidence is sufficient for that specific intended use, rather than accepting supplier assurance as a substitute for that judgement;
- independent verification where supplier evidence is incomplete, generic, or evaluated under conditions materially different from the customer's own.
MHRA Inspectorate guidance on AI-assisted content in regulatory submissions and inspection responses makes the same point from a different angle: using an AI tool, supplier-provided or otherwise, does not relieve the regulated organisation of responsibility for the accuracy, verifiability and technical review of what it submits. Accountability does not delegate to the tool or its vendor.
Retention and exit provisions
Contracts should preserve the customer's ability to retain and access evidence after the relationship ends, or after a specific model version is retired:
- retention of model version documentation and performance evidence for a period covering the customer's own record-retention obligations, not merely the supplier's commercial support window;
- a defined data-export mechanism and format on exit, including historical logs needed to support future investigations or inspections of records the AI touched;
- clarity on what happens to evidence access if the supplier is acquired, discontinues the product, or the contract ends for cause.
A supplier relationship that produces excellent evidence during the contract term but no contractual right to that evidence after termination leaves a customer unable to defend historical decisions once the commercial relationship that generated the evidence has ended: a foreseeable gap that belongs in procurement, not one discovered during a later audit.
Sources
- European Medicines Agency, Reflection Paper on the Use of Artificial Intelligence in the Lifecycle of Medicines (EMA/CHMP/CVMP/83833/2023, September 2024).
- U.S. Food and Drug Administration and European Medicines Agency, Guiding Principles for Good AI Practice in Drug Development (January 2026).