Your AI Agent Went Live in Four Weeks. What Exactly Went Live?
A four-week AI deployment can be completely real and still tell you very little about whether Quality should rely on the system.
The four-week claim is worth taking at face value. What's less certain is whether everyone in the room means the same thing by it.
An agent can retrieve approved procedures, follow configured instructions, draft a deviation classification and route it to a human reviewer within weeks. That's a legitimate technical achievement.
Three clocks, not one
I call this time to capability: when can the system perform the bounded task? It's only one of three clocks running underneath regulated AI.
- Time to capability: when can the system perform the bounded task?
- Time to confidence: when does the evidence justify the intended level of reliance?
- Time to maturity: when can the organisation change, monitor and improve the system without losing control of the evidence supporting that reliance?
These are clocks, not project phases. They can run at the same time and they don't move together. Capability can increase without confidence increasing alongside it, while a later model, prompt or retrieval change can preserve what users see while quietly changing the evidence behind reliance.
The state declaration below is what clock one actually produces: the capability line says what's true today, the evidence and reliance lines describe where clock two currently stands, and the residual line marks what remains outside the currently demonstrated state.
The state declaration
The practical response is to describe the state precisely: don't ask whether the system is live, ask for its state declaration. At the technical milestone, that means four lines documented. Capability: what can the agent currently do? Evidence: against which cases and conditions has that behaviour actually been demonstrated? Reliance: what are users permitted to rely on the output for, and where does human judgement remain authoritative? Residual: what remains explicitly outside the demonstrated state?
For a deviation agent, that might read: drafts classifications using approved procedures and selected historical records (capability); evaluated against a defined set of historical deviation categories (evidence); advisory output only, Quality retains classification and disposition authority (reliance); site-specific exceptions, novel event types and incomplete context remain outside the demonstrated state (residual).
Here's what that residual line catches in practice, call it the Legacy Precedent Trap. Consider a deviation agent that classifies a full historical set correctly except for one case, a closed deviation that, on inspection, was disposed of using a site practice Quality would no longer accept as good precedent. The agent didn't misread the rules it was given, it applied them exactly as written. What it got wrong was treating those rules as still-valid ground truth.
Historical Quality data contains approved decisions, retired interpretations, local workarounds and practices the organisation would not consciously endorse today. An agent can retrieve all of them perfectly and still produce the wrong organisational answer. This is one way AI can surface existing inspection debt: the system faithfully reproduces a decision the organisation never formally retired.
That's the actual question a state declaration exists to surface: what the organisation is currently willing to call correct, independent of whether the agent's retrieval was accurate.
That tells considerably more than "live in four weeks," and it gives Quality, Digital and the supplier the same description of what's actually been achieved.
What four weeks actually establishes
At week four, an agent may already retrieve defined information, follow configured instructions, produce the expected output, participate in a workflow, route that output through human review, and perform successfully against selected cases. Ask what that has actually established, and the honest answer is capability, not yet a justified reliance state. That's simply where the work stands at that point, not a mark against the implementation.
The pattern behind it isn't hypothetical. Suppliers in life sciences now publicly describe timelines of two to four weeks for initial deployments and four to six weeks for pilot-ready agentic applications, supplier-reported numbers rather than independent benchmarks, but consistent enough across the market to take the underlying claim seriously.
What none of those timelines describe is production in the sense a Quality organisation actually means it. A system can run in production without its output being relied on for a single GxP decision, and a controlled pilot can generate real evidence without ever being called production, deployment state and regulated reliance are simply different dimensions. The same slippage happens with the word live itself: it can mean the environment exists, that users can access it, that an integration works, that selected cases have passed, or that the organisation is genuinely relying on the result, several different claims hidden behind one word.
Go-live is an event. Reliance is a claim. The evidence required for the second doesn't appear merely because the first occurred, which is how two suppliers can promise the identical four-week timeline and leave a client in materially different states at the end of it.
The "pre-validated" problem
Supplier evidence can be extremely valuable. A vendor may have tested standard functionality, documented known behaviour, prepared reusable test assets and established controls before a client begins implementation, and that can reduce what the regulated company needs to generate from scratch. GAMP guidance takes a risk-based lifecycle approach to computerised systems that are fit for intended use, while ISPE's dedicated AI guidance extends that thinking specifically to AI-enabled GxP systems, including regulated-company and supplier responsibilities, monitoring, maintenance and continual improvement.
Pre-validation can reduce evidence-generation effort, though it can't pre-own the regulated company's own reliance decision. Rather than debating whether "pre-validated" is a fair marketing term, I'd rather ask directly: validated against what intended use, whose data, and whose Quality decisions? The answer tells you what the supplier evidence actually demonstrates and what remains for the regulated organisation to establish.
In practice, that's where the work actually starts: agreeing what operating state a vendor's milestone is actually meant to establish, before a validation protocol gets written — more on how Navata approaches AI governance and validation.
Speed is still the right call
Build the bounded capability quickly, put it in front of the people who understand the process, and challenge it in a controlled environment while failures are still cheap to investigate and correct. A system sitting in a slide deck generates little empirical evidence about how it will behave against the reality of a mature Quality process.
Rapid capability has real value because it brings the difficult questions forward. Which records should the agent consider? What constitutes acceptable evidence? Where should it stop? When is human judgement mandatory? Which conditions have never actually been tested? The faster the technical build becomes, the earlier those questions need owners.
Three questions before signing the SOW
A supplier proposing a four-week implementation wouldn't concern me simply because the timeline is short. I'd ask: what exact operating state exists at the end of week four? What evidence is included, and against whose intended use was it generated? What remains for us to establish before Quality relies on the output?
Those three questions make two apparently identical four-week proposals much easier to distinguish. One may deliver technical capability alone. The other may already have designed the path towards evidence, controlled reliance and operational ownership. The number on the proposal won't tell you which one you're buying.
The cost of getting it wrong
Confusing the three clocks fails in two different directions, and both are expensive. Treat capability as if it were reliance, and Quality ends up depending on evidence that was never actually established. Treat the whole deployment as a single clock, and the organisation delays the technical build waiting on confidence and maturity work that capability itself was never blocking. Regulated AI can be both fast and controlled. Managing three different clocks as though they were one is what breaks that.
The organisations that move fastest in regulated AI won't be the ones that collapse capability, confidence and maturity into a single validation programme. They'll be the ones that let capability move quickly, let reliance advance only as evidence actually justifies it, and build the organisational machinery to keep both under control as the system changes.
Four weeks may be enough to make the agent work. The strategic question is what has to become true before the organisation is entitled to depend on it.
Coming next
Part 2 — Time to Confidence
The Agent Works. Can Quality Trust It Yet?
The Legacy Precedent Trap in this piece cost the deviation agent one case: what the organisation had treated as ground truth, not how the model reasoned. Part 2 is about cases exactly like that one: what should happen between demonstrated capability and justified reliance, and how to tell whether a wrong answer is actually the AI's fault at all.
Part 3 — Time to Maturity
AI Doesn't Mature With Time. It Matures Through Organisational Intent.
Even after Quality has enough evidence to rely on the system, one question remains: when the system changes, how do you improve it without losing control of what's already been proven? That's where the third clock starts to matter.
More from Navata Insights → /insights
- ISPE, GAMP 5 Guide (Second Edition) — Appendix D11: Artificial Intelligence and Machine Learning
- ISPE, GAMP Guide: Artificial Intelligence (2025)
Views expressed are personal and do not represent any employer or client.