Navata
← All Library

Regulated AI

How to Assess and Govern a GxP AI Supplier

AI-assisted research and drafting · Practitioner reviewed by Rohith Karanam Sreedhar · 17 September 2026

Library content is researched and drafted with AI assistance and reviewed by a Navata practitioner before publication. For original analysis and long-form practitioner perspectives, visit Navata Insights →

An impressive model evaluation does not establish that a customer's workflow is fit for regulated use. The supplier may control the model, hosting and releases, while the customer controls prompts, retrieval sources, permissions, human review, permitted reliance and the final GxP action. Assurance depends on the joined service.

This page is a buyer and lifecycle reference. It does not determine EU AI Act classification, certify a supplier, or turn voluntary frameworks into legal requirements.

Define the use before assessing the supplier

Start with the intended-use statement: users, process, authorised inputs, exact task, output, permitted reliance, human decision, execution authority, operating conditions and prohibited uses. Classify the GxP risk separately from legal, privacy, information-security and enterprise risk classifications.

Decide which errors matter. An AI service that retrieves approved procedures creates different failure modes from one that drafts an investigation conclusion or writes fields into a quality record. Identify consequence, detectability, reversibility, exposure and dependence on the service.

Without this context, questionnaires produce a large collection of facts with no decision rule. Supplier evidence is useful only insofar as it supports a claim the customer needs to make.

Establish exactly what is being bought

Document the service boundary: product and model; hosted service, tenant and region; customer prompts, tools and retrieval; third-party models and subprocessors; operational data flows; retention and deletion; support access; availability and dependencies; and the change-release model.

Ask whether the supplier can substitute an underlying model or subprocessor without customer action. Determine which elements are shared across customers and which are isolated. Marketing product names often remain stable while the effective service changes underneath them.

Create an architecture and responsibility map showing who controls each component, record, decision and change. Identify where the customer can inspect evidence and where it depends on supplier attestation.

Examine data provenance and permitted use

Map prompts, documents, record fields, personal data, outputs, feedback, telemetry and support copies. Establish purpose, location, retention, encryption, access, deletion and whether data can be used for training or service improvement.

For retrieval-augmented generation, determine how sources are ingested, versioned, authorised, chunked, indexed, filtered and cited. Confirm whether deleted or superseded documents remain retrievable. For fine-tuning, identify the dataset, rights, preprocessing, exclusions, representativeness and lineage the supplier can disclose.

Do not assume that a model card describes the customer's data path. The complete system may include orchestration, retrieval, guardrails and workflow integrations supplied by different parties.

Evaluate quality and lifecycle controls

Seek evidence proportionate to intended use and supplier dependence. Review accountable roles; development, testing and release; versioning; configuration change; evaluation design and limitations; defect and incident management; security; privileged access; continuity and recovery; subcontractor oversight; customer notification; record retention; and end-of-service processes.

Ask for evidence, not only policy statements: sample change records, release notes, incident procedures, evaluation reports, independent audit scope, recovery-test summaries and known limitations where the supplier can appropriately provide them. Record evidence date, scope and limitations. A security report for one hosting environment or a certificate for one legal entity may not cover the purchased service.

Read performance evidence in context

Determine what task was evaluated, on which population, under which conditions, with what reference standard and by whom. Review sample size, class balance, language, rare cases, exclusions, repeated runs, variability and error taxonomy as relevant.

Ask whether evaluation data are independent of development data and resemble the customer's inputs. Aggregate accuracy can conceal unacceptable failure in a critical subgroup. A general question-answering benchmark does not establish controlled-document retrieval.

Separate model performance from system performance. The deployed outcome depends on prompts, retrieval, permissions, tool calls, post-processing and human review. Supplier evidence can reduce customer testing but rarely eliminates the need to challenge the configured end-to-end use.

Define evidence and record requirements

For each consequential output, decide what must be retained to reconstruct the event: user, timestamp, intended-use version, input or reference, source versions, instruction, model or service version, configuration, retrieved context, output, automated actions, human review and final decision.

Confirm which records can be exported, retention periods, identity and timestamp reliability, and behaviour after update or termination. Logging everything is not automatically appropriate; privacy and minimisation still apply. The objective is sufficient controlled evidence for the defined reliance.

Allocate responsibilities explicitly

The regulated customer retains responsibility for its GxP process and decision to rely on the service. The supplier may own product development, infrastructure and release evidence; the customer normally owns intended use, local risk assessment, configuration, authorised data, user access, procedures, training, end-to-end validation, acceptance and monitoring.

Supplier fitness does not establish action authority. The customer must separately decide which GxP actions the AI may propose, prepare, execute only with confirmation, or never perform autonomously. That authority decision should be explicit even where the supplier evidence, system validation and monitoring are otherwise strong.

Create a responsibility schedule for intended and prohibited use; legal classification; model, prompt, retrieval and workflow changes; access; performance monitoring; incidents; affected-record review; regulatory support; retention; deletion; and exit. Avoid “shared” without identifying who acts, decides, informs and retains evidence.

Put material controls into the contract

Procurement terms, data-processing terms, service levels and quality agreements should collectively address the identified risks. Depending on context, cover the defined service, authorised data use, hosting and subprocessors, security, assurance evidence, material-change notice, model substitution, incident notification, recovery, logs, retention, export, investigation support, suspension, termination assistance and verified deletion.

A contractual claim that a product is “GxP compliant” is not a validation strategy. Define observable supplier obligations that let the customer maintain its own controlled state.

Validate the configured use

Use supplier evidence where its relevance and reliability are understood. Add customer evidence for configuration, prompts, retrieval, roles, interfaces, data, procedures and permitted reliance. Test representative normal, boundary, adversarial and failure conditions.

For non-deterministic output, evaluate meaningful failure distributions rather than expecting exact repetition. Use challenge sets reflecting intended populations, verify provenance and assess whether reviewers can detect significant error. For automated actions, test authority boundaries, confirmation, rollback and prohibited actions.

Validate fallback and suspension. The process should remain controlled when the service is unavailable, degraded or deliberately disabled. Confirm backlog, later reconciliation and whether users may bypass safeguards under pressure.

Document exclusions. Performance in one language, document family, site or model version should not authorise broader use silently.

Govern supplier and service change

Define material-change categories before go-live: underlying model, training or fine-tuning, prompt and orchestration, retrieval corpus, safety filters, hosting, subprocessor, retention, API, performance, output format and customer workflow.

Require notice sufficient for impact assessment. Not every release requires full revalidation, but every applicable change needs a disposition. The customer should be able to hold, test, restrict or suspend use where evidence is insufficient.

Maintain a service baseline identifying components and conditions supported by current evidence. Compare each change against intended use, risks and prior tests. Re-establish affected evidence and record why the remainder remains applicable.

Monitor operational use

Monitor signals connected to actual risk: source-citation failures, unsupported statements, reviewer overrides, prohibited-action attempts, subgroup performance, retrieval gaps, latency, unavailable dependencies, unusual access, incidents and workarounds. Define thresholds, cadence, investigation and disposition.

Supplier platform metrics may not reveal local harm. The customer needs feedback from final decisions and records, while avoiding monitoring that captures only errors users notice. Periodically reassess intended use, service boundary, supplier status, performance, incidents, changes, access, procedures and unresolved limitations.

FDA and EMA's January 2026 Good AI Practice principles for drug development support clear context of use and risk-proportionate lifecycle controls. NIST AI RMF offers a voluntary structure for governing, mapping, measuring and managing AI risk. These sources inform the method; neither is a universal pharmaceutical validation checklist.

Respond to incidents and corrections

Agree how the supplier identifies affected versions, tenants, dates and outputs. Translate that technical scope into affected GxP records and decisions. Restoring service does not resolve potentially incorrect historical outputs.

Preserve the notice, timeline, affected population, containment, record review, corrections, root cause, corrective action and evidence re-established. Decide when to disable the service and how to communicate limitations.

Design exit before dependence develops

Define export of inputs, outputs, configurations, prompts, retrieval sources, logs, evaluation records and administrative history as applicable. Confirm formats, fees, timescales, assistance, deletion and post-termination access.

Assess portability realistically. A replacement model may not reproduce the same outputs or evidence. Identify the fallback process, backlog control and new-supplier validation. Avoid architectures in which essential regulated records exist only inside a supplier interface.

Use staged decisions

Apply explicit gates: intended use and risk; service and data boundary; supplier evidence and contractual controls; end-to-end validation; operational readiness; lifecycle monitoring; and exit. Each gate needs an accountable decision-maker, required evidence, open issues and a permitted disposition: proceed, proceed with conditions, restrict, remediate, suspend or reject.

Important boundaries

The EU AI Act establishes legal obligations where applicable; classification and actor roles require legal analysis. NIST AI RMF is voluntary. FDA and EMA Good AI Practice principles are broad lifecycle principles for drug development and do not make every practice a binding requirement in every context. Security certifications have defined scopes and do not establish GxP fitness.

The staged decision model, evidence questions and contract-control examples are Navata Library methods and must be tailored.

Sources