The Most Important AI Control May Be the Action It Cannot Take
A human approval step proves little once the agent has already shaped the evidence, prepared the decision and holds the permission to perform the regulated act.
A deviation investigation reaches five people before it closes.
The investigator drafts the proposed conclusion. The investigation owner approves it. Quality reviews the record. The CAPA owner accepts the actions. A final approver confirms closure.
The approval history shows five names.
Each of those five people read the same AI-generated summary. The same agent selected the evidence, reconstructed the chronology, drafted the root-cause narrative, and recommended the actions. The source records were available, a couple of clicks away, but the workflow surfaced the summary first. When a reviewer asked a follow-up question, the same agent answered it, using the same retrieval logic that produced the original recommendation.
Did five people independently review the investigation? Or did five people approve one machine-generated interpretation, five times over?
I think most conversations about human oversight stop at the first question: did a person approve the output. They rarely ask the second, harder one: what authority did the system already exercise before that approval task ever appeared on a screen.
A human in the workflow doesn't, by itself, establish human control. I want to know what the agent can read, what it can select, what it can prepare, what it can change, and which actions stay technically out of its reach regardless of who's watching.
The record shows five approvals
A long approval history can look like control. It shows that authorised people participated, and it can show the sequence, the timestamps, the comments, the electronic signatures. What it doesn't show is whether those people reached their conclusions through sufficiently independent review.
I separate two ideas that get treated as the same thing. Approval multiplicity is several people clicking approve. Analytical independence is each accountable reviewer having real access to the evidence and a genuine basis for their own judgement, not a rerun of someone else's.
I don't expect every approver to repeat the whole investigation, that adds effort without adding much control. I do expect the workflow to preserve the actual purpose of each approval, which means each role needs a different slice of the record to do its own job properly, not the same generated account everyone else received.
When every role receives that same generated summary instead, their different responsibilities collapse into a single analytical path. The investigator assumes Quality will check the source. Quality assumes the investigator already verified the reconstruction. The final approver relies on the fact that everyone before them already signed. Five names end up representing five acts of confirmation around one machine-framed account, and the weakness in that isn't visible from the approval count.
Nobody in that chain has acted irresponsibly. The workflow has simply let each person rely on somebody else, while the same AI interpretation sits underneath every one of the five approvals.
Start with authority, not accuracy
Most AI controls start with the output. Was the answer correct. Did the model hallucinate. Did a human check it. Those are fair questions, and they don't cover the risk once a system can act inside a regulated workflow.
I split an agent's authority into three layers. Interpretation, where it reads records and produces an assessment. Recommendation, where it proposes a classification, a root cause, a CAPA, an owner. Execution, where it changes the state of the process: updating a field, routing a record, closing an action, making a document effective.
Interpretation and recommendation create information risk, something can be wrong and someone can catch it. Execution creates a regulated consequence the moment it happens. An accurate recommendation doesn't automatically justify execution authority; a model can perform exactly as intended and still have no defensible reason to hold the permission that closes the record.
The FDA and EMA's joint Guiding Principles of Good AI Practice in Drug Development, issued in January 2026, set human oversight and risk-based lifecycle management as baseline expectations for AI across the product lifecycle. What neither agency spells out is how that oversight should actually be architected inside a workflow. That's the gap this piece is written into.
Veeva and platforms like it make workflow actions easy to describe technically. A record moves from one state to another. A field value changes. A task completes. A document becomes effective. The transaction is simple. The meaning isn't.
I looked at how to validate whether an AI's answer holds up in How to OQ/PQ a Hallucination. That question is about whether the output performs within its intended use. The question here comes before it: even when the output clears that bar, should the agent have had the authority to act on it at all?
Some actions are more than transactions
Some GxP actions represent an attestation rather than a transaction, a named person reviewed the evidence, exercised judgement, and accepted responsibility for what happens next. Applying an electronic signature. Approving or rejecting a regulated record. Assigning a final root cause. Confirming CAPA effectiveness. Accepting a residual quality risk. Approving product disposition. Closing an investigation. Making a controlled document effective.
I'd start from prohibition on all of these and require a positive, documented reason before granting an agent execution authority over any of them. A platform being technically capable of it isn't a reason.
That doesn't mean the agent stays uninvolved. Take CAPA effectiveness. An agent can analyse recurrence data, retrieve related deviations, and draft an effectiveness assessment, genuinely useful, faster, more consistent. Confirming effectiveness is a different act. It asserts that the defined evidence was considered, the criteria were met, and the CAPA achieved its intended outcome. Letting the agent recommend effectiveness doesn't mean letting it confirm effectiveness, and I'd keep those as two separate, deliberately made decisions rather than one setting that quietly covers both.
The agent can shape the decision without approving it
An agent doesn't need approval authority to shape what gets decided. It only needs to control what the reviewer sees: which records get retrieved, which count as relevant, how the chronology gets built, which contradictions surface, which exceptions get compressed into a line, which recommendation appears first.
Picture an investigation with ten records pointing towards operator error, two suggesting a recurring equipment fault, and one maintenance record that contradicts the proposed sequence entirely. The agent retrieves the operator records and drafts a clean summary. It leaves the maintenance record out because it scored below the retrieval threshold. Every sentence in that summary can be factually correct. The conclusion can still be wrong, because the evidence window it was built from was incomplete.
This kind of failure is harder to catch than a hallucination. A fabricated fact can be challenged the moment someone checks it against the source. An omission leaves nothing on the screen to challenge: the summary reads cleanly, the chronology looks complete, the recommendation looks supported, and the missing record stays invisible unless a reviewer goes looking for something they don't know is missing.
Accuracy within a selected evidence window doesn't prove the window was selected correctly.
One AI can sit behind every approver
Shared evidence is normal, several people reviewing the same source material is exactly what should happen. Shared machine framing is a different condition: several people relying on the same AI interpretation as the primary basis for their review. That's where the agent becomes the hidden analytical dependency running underneath an approval chain that looks, from the outside, like five separate judgements.
I'd watch for the design features that make passive confirmation likely: the summary shown up front while source evidence sits behind extra clicks, approval taking one click while rejection needs a written justification, the AI's recommendation pre-populated into the controlled fields, later approvers seeing earlier approvals before they've looked at the evidence themselves, the same agent answering every challenge with the same evidence set and ranking logic that produced the original answer.
Training reviewers to stay vigilant doesn't fix any of this. It doesn't change the review surface or the incentives the workflow itself creates.
What does help is not giving every role the same screen. Route the investigator to the full chronology and source evidence, Quality to the contradictions and procedural exceptions, the CAPA owner to the causal rationale and its weak points, the final approver to the closure criteria and whatever risk is still unresolved. Everyone can still reach the complete record. The default view should be built around what each role is actually accountable for catching, not repeat the same account five times over.
Reversible in software, irreversible in practice
An administrator being able to reopen a closed record doesn't undo what happened while it was closed. I separate technical reversibility, whether the platform can restore the earlier state, from regulated reversibility, whether every material consequence created by the action can also be identified and corrected. The second is the harder test, and the one that actually matters.
An SOP made effective has already driven training. A deviation closed and dropped from the overdue report has already stopped getting attention. A classification that avoided escalation has already missed its window. A CAPA declared effective has already ended the monitoring that would have caught a recurrence. The system transaction can be reversed. What happened, and what didn't happen, while it stood, usually can't be.
The practical rule: the harder an action is to undo inside the regulated process itself, not the database, the less execution authority the agent should hold over it.
The Non-Delegable Action Test
A matrix sorting actions into low, medium, and high risk isn't specific enough for this. I use four questions instead.
- State. Does the action change the regulated state of a record, process, or product? Draft to effective, open to closed, proposed to approved, under investigation to complete, pending disposition to released. A field can change regulated meaning even when the workflow status stays the same.
- Attestation. Does it assert that evidence was reviewed, judgement was exercised, or responsibility was accepted? An approval, a signature, an effectiveness confirmation, these are attestations, not workflow clicks, and assigning a root cause or accepting a residual risk can carry the same weight even when the system only logs it as an ordinary field update.
- Evidence control. Can the actor determine which evidence the next decision-maker sees? An agent can hold real authority through retrieval and ranking alone, without ever touching an approval button.
- Consequence. Can the action create an effect that reversing the system transaction won't fully remove? Training already delivered, reports already filed, escalation windows already missed.
A yes on any of these shifts the burden of proof onto the exception. I'd start from the action being technically unavailable to the agent, and treat any exception as something that earns its own documented case: intended use, business justification, authorised scope, the accountable human role, the evidence the human actually had, the confirmation step, how it gets reversed, what it affects downstream, how it gets monitored once it's live, and validation of both the approval path and the rejection path.
Four boundaries, not one approval step
Translated into architecture, this becomes four boundaries rather than a single review gate. The proposal boundary covers what the agent may draft or recommend, no restriction beyond normal review. The preparation boundary covers what it may assemble without completing, populating proposed values or placing a transaction in a queue, without that queue entry counting as a decision. The execution boundary covers what it may do only once a named, authorised person confirms that specific transaction, not a broad permission granted once when the agent was switched on and left standing. The prohibition boundary covers what it should never be able to do at all: apply a human electronic signature, approve its own recommendation, alter the evidence set once formal review has started, close a regulated investigation, confirm CAPA effectiveness, make a controlled document effective.
The prohibition boundary is the one that matters most. It has to be enforced in permissions, workflow design, and test evidence, somewhere a policy document can't reach. A policy saying the agent “must not close records” is weaker than an account that lacks the permission to close them.
I test whether the human can genuinely disagree
None of this holds up if the person on the other end can't actually exercise judgement. That means reaching the source evidence, seeing what the agent chose not to surface, telling generated content apart from source content, changing or rejecting the recommendation outright, and stopping the transaction before it executes rather than after. It also means the interface stops treating the AI's answer as the default the moment a reviewer replaces it, rather than quietly keeping it anchored in view.
It's worth testing the rejection path specifically, not just the approval path. How many steps does approval take against how many steps rejection takes. Does rejection drop the reviewer into an unstructured manual process. A workflow where approval is one click and rejection is a written justification satisfies the letter of human review while quietly training people towards the click. I'd build rejection, correction, and independent-review scenarios into validation itself; testing only the expected approval path proves the workflow can process agreement, and says very little about whether human authority still means anything when the person disagrees.
The strongest control may be a missing permission
The usual answer to AI risk is another human approval. That's fine for drafts and for low-impact recommendations. It stops being enough once the agent has already determined the evidence, set the proposed values, shaped every reviewer's interpretation, and holds the permission to perform the final act itself.
Five signatures don't add up to five independent judgements when every reviewer is following the same machine-generated path. A reversible field change doesn't remove the consequences already created in the regulated process. A policy instruction doesn't stop execution when the agent still holds the permission to execute.
For each proposed use, I'd ask one question: does this action become sufficiently controlled through human review, or does it become controlled only once the agent can't perform it at all.
For some GxP decisions, the strongest control is not another approval step. It is the absence of the agent's permission to act.
If you're working through where these boundaries should sit in a live Vault or QMS build, that's the conversation I have with clients at Navata. More on how I approach AI governance and validation at /ai.
Views expressed are personal and do not represent any employer or client.