Navata
← All insights
Quality AI August 2026 · 8 min read
The Three Timelines of Regulated AI · Part 3 of 3 · Time to Maturity

AI Doesn't Mature With Time. It Matures Through Quality Intent.

Maturity, for a regulated AI system, is the organisation's capacity to correct behaviour it no longer trusts, and prove afterwards that its Quality Intent, its architecture and its evidence are still describing the same system.

I'm calling the mechanism behind that the Correction Path, the third and final framework in this series.

Part 1 separated capability from reliance, and introduced the three clocks, time to capability, time to confidence, time to maturity, along with the Legacy Precedent Trap: a deviation agent that classifies a closed historical case correctly, by the rules it was given, and still gets it wrong, because the case was disposed of using a site practice Quality no longer accepts as good precedent. Part 2, The Agent Works. Can Quality Trust It Yet?, showed how Quality earns confidence, and diagnosed failure through the Navata Quality AI Failure Map, six categories, five inside the AI system, one, Quality Intent, outside it, plus the Correction Record that captures what each correction found.

Someone corrects the agent. A Correction Record gets written up properly against those five fields. The record is dated, specific, and for a moment the problem feels closed.

It describes what one person caught and fixed. Whether the system's operating state, what determines its behaviour now, has moved to match it is a separate question, and it stays open until something checks the operating state against the evidence state: the qualification record that's supposed to justify relying on the system in the first place.

I've seen this pattern play out. A reviewer catches a bad answer and documents it properly. The fix goes in at the prompt or retrieval layer within a week. Everyone moves on. Weeks later a different reviewer catches almost the same failure in a different record, and nobody connects it to the first, because nothing formally links them. The operating state moved. The evidence state, and everything built on it, stayed exactly where it was.

Diagram: Operating state vs evidence state. Operating state (what determines its behaviour now) and evidence state (the qualification record that's supposed to justify relying on the system in the first place) run level until a correction is applied. After the correction, the operating state rises to a new level while the evidence state stays flat, opening a gap: behaviour has changed but the evidence that justifies reliance has not. The operating state moved. The evidence state, and everything built on it, stayed exactly where it was.

Classify before you fix

A correction is material when it could change what evidence the system treats as persuasive, what conclusion counts as acceptable, or when the process should escalate. Technical size doesn't predict this either way. A one-line instruction edit can matter more than a substantial configuration change, because that line decides what the agent does with ambiguity next time.

Two examples show the difference. A reviewer catches the agent quoting a superseded SOP because a retirement flag never propagated through the source corpus. Fixed once in the source layer, it doesn't recur, and nothing about what counts as acceptable evidence changes, so it isn't material. A reviewer catches the agent framing a borderline deviation as procedural rather than product-impacting. Fix that one classification and the next borderline case can fail the same way, because what the system treats as an acceptable conclusion hasn't actually moved, so it is.

The Legacy Precedent Trap is the clearest version of this I have. Nothing about it looks like a defect. The agent read the record correctly, reasoned from it correctly, and still produced an answer Quality wouldn't stand behind. Judged by technical size, there's no correction to make at all. Judged against the materiality test, it's one of the largest kinds there is.


Attribute before you configure

Once a correction reads as material, the next question is which layer it actually belongs to, because that decides what "fixed" has to mean. I use the same six categories from the Failure Map here, because attribution and diagnosis are the same exercise viewed from opposite ends.

A source-layer correction is usually content governance operating under an AI label. It gets fixed where content already gets fixed, and rarely touches the system's validated configuration at all. An instruction or reasoning-layer correction is different in kind. It means the system's actual behaviour has to change, and that's the category most change-control processes were never built to see, because nothing about it resembles a code deployment or a workflow edit.

Quality Intent corrections are among the hardest to place. The Legacy Precedent Trap belongs here. The system was configured correctly and did exactly what it was told to do. What Quality means by acceptable precedent had simply moved on since that configuration was written, and nobody had captured the new version anywhere the system could be checked against. There is nothing for a validation script to test until someone writes down what changed.

If you're working out where a correction like that should sit in your own change-control architecture, that's the design work I do with clients building AI governance into a live Vault or QMS. More at /ai.


Propagate before you close the record

Attribution tells you where a correction belongs inside the AI system. Propagation is the separate question of where the same decision needs to exist outside it, and it's the step most programmes skip, because nothing in a standard change-control form prompts for it.

Suppose the Legacy Precedent Trap correction plays out fully. Quality decides the old site practice no longer counts as acceptable precedent, and the agent's instruction gets updated to reflect that. The immediate behaviour is fixed. But the SOP governing deviation classification may still describe the old practice as an acceptable reference point. The SME who trains new investigators may still be teaching the older interpretation, because nobody told them it changed. A second AI system, drawing on the same historical investigation records for a different use case, has no idea the precedent was ever revised. And the original challenge set used to qualify the agent may still score the old, now-unacceptable answer as correct.

Every material correction has a propagation boundary: the procedures, guidance, related use cases and test material built on the assumption the correction just overturned. Checking that boundary means confirming, item by item, what needs to change, and recording the answer either way, even when the answer is nothing. Skip the check and the same organisational decision ends up living in several versions at once, each one confidently wrong about the others.


Where the change actually lives

This is closer to established territory than it first looks. Expanding on GAMP 5 Second Edition's Appendix D11, the July 2025 ISPE GAMP Guide: Artificial Intelligence addresses dynamic systems, ones whose behaviour can change through retraining, tuning or drift, across the concept, project and operational phases of their life cycle. Separately, the Guide covers change management addressing both business and quality drivers, and CAPA considerations specific to AI-enabled systems, alongside guidance on using live production data for ongoing monitoring.

Read together, that's a more useful anchor than it sounds. A system whose behaviour can legitimately change without a formal release event, combined with guidance that already expects change management and CAPA to cover AI-enabled systems, means a correction with no configuration line to point to still sits inside what the framework anticipates. Most change-control processes just aren't built to route it there yet.


Re-anchor the evidence state

A material correction, once attributed and propagated, has one more step before the record can close: it has to reach back and check the evidence state that justified the system's current level of reliance in the first place.

That evidence was built against a defined boundary of scenarios and failure modes covered when reliance was justified. A correction that reveals a failure mode inside that boundary, the way the Legacy Precedent Trap did, means the evidence state is now incomplete too, alongside the operating state. Fixing the system without updating the evidence leaves Quality relying on something that has genuinely improved, backed by a qualification record that no longer describes it accurately. That gap stays invisible until an inspector asks for the evidence behind a reliance decision, and the trail stops one step short of the correction that actually produced today's behaviour.

Re-anchoring scopes to what the correction actually disturbed. It means working out, deliberately, which parts of the existing evidence still hold, which need supplementing, and whether the reliance boundary Quality signed off on is still the right one. Most corrections touch a narrow slice of that evidence rather than all of it.

Diagram: The Navata Correction Path. 1 Classify — materiality test: could this change what evidence counts, what's acceptable, or when to escalate? 2 Attribute — which layer does it belong to: source, retrieval, context, instruction, reasoning, or Quality Intent? 3 Propagate — where else does this decision need to exist: procedures, guidance, related use cases, test material? 4 Re-anchor — does the evidence that justified reliance still hold, or does it need updating too?

Digital owns the mechanism

Classification, attribution, propagation and re-anchoring are architecture. They need a defined owner, a defined workflow and an audit trail, like any other change-control process, and that ownership sits with Digital. A mechanism nobody is accountable for maintaining degrades back into the pattern this piece opened with.

The judgement underneath it belongs somewhere else. Whether a correction still matches what Quality intended when reliance was first granted, or is quietly redefining what "correct" means, has to stay a call for whoever holds Quality Intent, made explicit and evidenced every time, rather than assumed from the fact that a fix shipped and nobody complained.

The first two parts of this series asked whether the agent works, and whether Quality could trust it. This one asks whether that trust survives the correction that was always going to come. The Correction Path is the answer I use: classify what the correction actually is, attribute it to the layer it belongs to, propagate the decision everywhere it needs to exist, and re-anchor the evidence reliance was built on. Skip any one of those steps and a caught mistake becomes paperwork instead of a change. Run all four, and the system keeps changing without ever losing the thread back to why it was trusted in the first place.


The full series

The Three Timelines of Regulated AI
1 Time to Capability Your AI Agent Went Live in Four Weeks. What Exactly Went Live? Read → 2 Time to Confidence The Agent Works. Can Quality Trust It Yet? Read →
3 Time to Maturity AI Doesn't Mature With Time. It Matures Through Quality Intent. Reading now

More from Navata Insights → /insights

Sources
Regulatory note

The Navata Quality AI Failure Map, Quality Intent, the Correction Path and the Three Timelines of Regulated AI are practitioner frameworks developed by Navata. They are not regulatory classifications and do not replace applicable regulation, guidance, Quality Risk Management, validation or organisation-specific procedures.

Views expressed are personal and do not represent any employer or client.

About the author
Rohith Karanam Sreedhar
Founder & Principal
Navata

Navigate the Complex. Architect the Compliant.