Navata
← All Library

Regulated AI

How to Classify the GxP Risk of an AI Use Case

AI-assisted research and drafting · Practitioner reviewed by Rohith Karanam Sreedhar · 8 September 2026

Library content is researched and drafted with AI assistance and reviewed by a Navata practitioner before publication. For original analysis and long-form practitioner perspectives, visit Navata Insights →

“Uses generative AI” is not a risk classification. Neither is “human in the loop”. Two applications of the same model can need materially different assurance where one helps search approved procedures and the other prepares or executes a quality decision.

Classification should begin only after intended use is specific enough to identify the user, task, inputs, output, permitted reliance and prohibited actions. Without that boundary, the risk assessment is rating a technology label rather than a regulated use.

Keep three classification systems separate

A GxP assurance tier, an enterprise information-security rating and a legal classification answer different questions. They may inform one another, but they should not be merged casually.

The EU AI Act establishes statutory categories and obligations where the Regulation applies. Whether a use is prohibited, high-risk or subject to another obligation is a legal assessment. A use outside an EU AI Act high-risk category may still carry substantial GxP risk because a quality decision relies on it. Conversely, a legal classification does not prescribe the complete validation approach for a particular pharmaceutical process.

The tiers below are Navata planning categories. They are not regulatory labels.

Assess six dimensions

1. Regulated consequence

Identify the most material plausible result of an undetected wrong, incomplete or unavailable output. Consider effects on patient protection, product quality, data integrity, regulatory commitments and the ability to reconstruct a decision.

Do not rate only direct harm. An incorrect summary may delay an investigation or hide recurring events. A classification recommendation may determine escalation. A retrieval omission may deprive an approver of evidence without changing any field itself.

2. Permitted reliance

State what users may do because of the output:

  • use it only for navigation;
  • use it as a draft that must be independently reconstructed;
  • consider it as one decision input;
  • adopt it after defined verification;
  • allow it to populate controlled fields;
  • allow it to trigger or execute workflow action.

The same output becomes higher risk as permitted reliance increases. Training users to be cautious does not compensate for an interface and procedure that treat the answer as authoritative.

3. System authority

Separate interpretation, recommendation, preparation and execution. An AI can shape evidence and prepopulate a decision without applying the final signature. Record what it can read, select, write, route, approve or close, and which actions are technically unavailable.

Where the action changes regulated state, represents an attestation or creates a consequence difficult to reverse, require a specific justification for system authority rather than assuming human review is sufficient.

4. Detectability before consequence

Ask whether an informed reviewer can detect the material failure before reliance. This depends on the review surface and evidence, not the presence of an approval box.

Detection is stronger when the reviewer can:

  • reach cited source material;
  • see omissions or retrieval limitations;
  • distinguish generated from source content;
  • compare against defined criteria;
  • reject or correct without disproportionate effort;
  • stop execution before state changes.

Detection is weaker where errors are plausible, fluent, distributed across many records or visible only after downstream consequence.

5. Reversibility and containment

Technical reversibility is not always regulated reversibility. Reopening a record may not recover a missed escalation or undo training delivered from an incorrectly effective procedure.

Consider how quickly the error can be contained, how the affected population can be identified and whether every downstream action can be corrected. Weak traceability increases risk because even a reversible action may have an unknowable impact boundary.

6. Exposure and variability

Assess frequency, scale, data diversity and the likelihood that operating conditions depart from validation. A low-consequence task performed millions of times can create material cumulative exposure. A narrow high-consequence use may remain controllable through strong prohibition and review boundaries.

Include model, prompt, retrieval, reference-data, interface and supplier variability where each can change outputs.

Use tiers to route assurance work

A practical three-tier model can work if its consequences are defined.

Tier 1: bounded assistance

The output supports navigation or drafting, carries low regulated consequence, cannot execute actions and is readily checked before use. Typical assurance may include intended-use approval, source and access controls, representative functional testing, basic monitoring and clear user instruction.

“Draft” does not automatically mean Tier 1. A draft displayed as the default account to every approver may shape a consequential decision.

Tier 2: consequential decision support

The output informs a regulated judgement or populates proposed controlled values, but an authorised person verifies defined evidence and confirms the action. Assurance normally needs explicit performance properties, representative and adverse cases, retrieval or data validation, review-interface testing, defined override and rejection, traceability, change assessment and monitored failure signals.

Tier 3: high-consequence or execution-capable use

The AI can perform or materially determine a regulated action, errors are difficult to detect before consequence, effects are hard to reverse, or the affected population cannot be contained reliably. Tier 3 should trigger senior Quality and architecture review, strong technical authority boundaries, expanded validation, independent challenge, continuity planning and a presumption that certain actions remain non-delegable.

Some uses should be prohibited rather than assigned more controls. A tiering process should preserve that outcome.

Apply a reproducible classification sequence

Use the same sequence for every intake so a familiar supplier or enthusiastic sponsor cannot bypass the difficult questions.

  1. Confirm intended use. Reject classification if the user, task, inputs, output, permitted reliance or prohibited actions remain undefined.
  2. Identify the regulated consequence. Describe what can happen if the output is wrong, incomplete, biased, unavailable or used outside scope.
  3. Map reliance and authority. Record what the user may rely on and what the system can select, prepare, write, route or execute.
  4. Assess pre-consequence detection. Test whether an authorised, informed person can recognise the material failure before action.
  5. Assess containment and reversibility. Determine whether affected records and downstream actions can be identified and corrected.
  6. Assess exposure and variability. Consider volume, frequency, populations, data diversity and components capable of changing behaviour.
  7. Apply prohibition and escalation rules. Decide whether the use is acceptable for tiering at all.
  8. Assign the tier and evidence path. Record the rationale, owner, approver and residual uncertainty.

If evidence for a dimension is absent, classify conservatively or return the use for clarification. Do not convert “unknown” into a low score.

Use explicit tier-decision rules

Assign Tier 3 when any of these conditions applies unless the intended use or technical architecture is redesigned so that the condition no longer exists, and that changed boundary is demonstrably enforced:

  • the AI can execute a regulated state change or attestation with material consequence;
  • a material error is unlikely to be detected before reliance;
  • consequences cannot be contained or reversed reliably;
  • the affected population cannot be reconstructed;
  • the output materially determines product disposition, patient protection or acceptance of significant quality risk.

Assign Tier 2 when the output informs a consequential judgement or prepares controlled values, but a qualified person can examine defined evidence, reject the output and stop action before consequence. Tier 2 also applies where individual consequences are moderate but scale or repeated reliance creates material exposure.

Assign Tier 1 only when reliance is bounded, consequences are low, the output is readily checked, the system cannot perform regulated action and failure remains visible and recoverable. Every Tier 1 condition must remain true; a single material exception moves the use upward.

The final tier is not an average. Strong detectability should not mathematically cancel an irreversible high consequence. Use the highest material dimension as the starting point and lower it only where a specific, tested control changes the intended use itself.

Resolve conflicting dimensions

Conflicting ratings are expected. Use these rules:

  • High consequence, strong detection: retain the higher tier unless detection occurs before action, uses sufficiently independent evidence and has been tested under realistic workload.
  • Low consequence, very high scale: consider Tier 2 where cumulative omissions or systematic bias could create a material population effect.
  • High authority, apparently reversible action: assess regulated reversibility, not the ability to undo a database transaction.
  • Low authority, poor detectability: increase the tier where machine framing can quietly shape a decision even without execution permission.
  • Stable model, changing retrieval or workflow: classify the whole use; model stability does not remove system-level variability.

Document the conflict and why one dimension governs. This is more defensible than a single composite score whose arithmetic hides the critical condition.

Worked classification examples

Bounded assistance

An assistant searches approved procedures and returns cited passages to a trained user. It cannot update records. The user opens the cited procedure before relying on it, and missing or unavailable sources are visible. Consequence is bounded, authority is absent, detection is strong and the action is reversible. This can be Tier 1, subject to retrieval, access, citation and monitoring controls.

If the interface begins answering without citations and the SOP permits users to follow the answer directly, permitted reliance changes. The same model and search task may then require Tier 2.

Consequential decision support

An AI reviews deviation records and proposes a severity classification. An investigator can see the source evidence, change the value and must provide the final decision before workflow continues. The output shapes escalation and due dates, so consequence and reliance are material even though execution remains human-confirmed. This is Tier 2 and needs representative and adverse cases, review-interface evidence, override testing, traceability and monitored disagreement.

Execution-capable AI

An agent evaluates an investigation package, writes final root cause and closes the record when confidence exceeds a threshold. Closure is a regulated state change, an error may end escalation and later reopening cannot undo missed action. This meets Tier 3 and may also meet prohibited-use criteria. Adding a periodic sample review after closure does not create pre-consequence detection.

Define the assurance and approval path

Control areaTier 1Tier 2Tier 3
Intended useBounded task and prohibited actionsDetailed reliance and human-decision boundaryFull authority, consequence and non-delegable-action analysis
Validation evidenceRepresentative functional and retrieval checksDefined performance properties, adverse cases and end-to-end workflow testsExpanded independent challenge, execution-boundary and containment evidence
Human oversightUser can verify source before useRole-specific evidence view, correction, rejection and stop authority testedIndependent approval architecture; no reliance on after-the-fact sampling
PermissionsNo regulated executionProposed values or preparation separated from confirmed actionLeast privilege, technical prohibitions and transaction-specific authorisation
MonitoringAvailability, retrieval and material user feedbackPerformance, overrides, omissions, data and workflow changesContinuous high-risk indicators, rapid containment and senior escalation
Change controlReassess intended-use or source changesAssess model, prompt, retrieval, workflow, permission and reliance changesFormal Quality/architecture decision before material change; revalidation where conclusions weaken
Approval routeProduct and process owner under defined governanceProcess owner, Quality and validation/AI assuranceSenior Quality plus accountable architecture and business authority

Tier 2 should escalate where the reviewer lacks necessary evidence, rejection is not operationally viable, or monitoring shows systematic reliance on the recommendation. Tier 3 should escalate whenever proposed controls depend on policy while technical execution authority remains available.

Define prohibited-use criteria

Return or prohibit a proposed use when:

  • it would impersonate or apply a named person’s electronic signature;
  • it would approve its own recommendation or evidence selection;
  • the evidence needed to detect a material error cannot be made available before action;
  • the affected population cannot be identified after failure;
  • the organisation cannot retain or reconstruct evidence supporting consequential output;
  • required human responsibility would exist only nominally because the person cannot intervene;
  • applicable law, regulation or policy prohibits the use.

Prohibition can be narrowed by redesigning intended use—for example, removing execution permission and retaining cited decision support. It should not be bypassed by assigning the highest tier without changing the unacceptable condition.

Translate classification into an evidence plan

For each tier, define minimum expectations for:

  • accountable owner and approver;
  • intended-use detail;
  • supplier and model information;
  • data and retrieval evidence;
  • performance and failure-condition testing;
  • human-oversight design;
  • permissions and execution limits;
  • record and audit evidence;
  • monitoring and periodic review;
  • change triggers and reassessment;
  • fallback and retirement.

FDA and EMA's January 2026 Good AI Practice principles, within their drug-development and product-lifecycle scope, support clear context of use and risk-proportionate validation, mitigation and oversight. They inform the dimensions used here but do not prescribe Navata's tiers or a universal GxP AI classification method. NIST AI RMF similarly maps risk through context, impacts, human oversight and lifecycle management.

ICH Q9(R1) provides a general quality-risk foundation: decisions should be science-based and effort, formality and documentation should be proportionate. The final classification should therefore retain the rationale and evidence behind each rating rather than only a colour or score.

Reassess when the reliance changes

Trigger reassessment when there is a material change to:

  • intended task or user population;
  • permitted reliance or execution authority;
  • model, prompt or orchestration;
  • data sources, retrieval corpus or metadata;
  • workflow, permissions or review interface;
  • supplier hosting, logging or retention;
  • observed performance or failure pattern;
  • ability to detect, contain or reverse error.

A supplier may describe a model update as minor while the customer's risk changes because retrieval, latency or output format affects the review path. The system owner should classify the effect on the use, not inherit the supplier's release label.

For example, a Tier 1 procedure-search assistant is changed to prepopulate a deviation’s proposed procedural-breach field. The model is unchanged and citations remain available, but the output now enters a controlled record and anchors the investigator’s classification. Reassessment moves the use to Tier 2, adds representative deviation cases, reviewer correction and rejection tests, traceability from proposed value to sources, permission controls and monitored override behaviour. The trigger is changed reliance and workflow authority, not a model release.

Keep classification, validation and acceptance separate

Initial classification routes the use into a proportionate assurance and approval path using the best available design information. It does not prove performance.

Validation strategy defines the claims, test methods, data, conditions and acceptance evidence needed for that classified use. Successful validation supports reliance within the approved boundary; it does not decide whether residual risk is acceptable.

Ongoing risk acceptance is the accountable decision to operate with known residual limitations, monitoring and change controls. It must be revisited when evidence or conditions change. Keeping these decisions separate prevents a high-level risk tier from being treated as validation evidence or a successful test report from silently accepting a governance risk.

Sources