Skip to content
Pyrana
Resources
AI and data leadersGovernanceRegulated industriesAgentic applications

Trust, but Validate

Nobody in a regulated industry asks whether they trust their ERP. They ask whether it is validated, who has access, and where the audit trail is. Here are the eight things it takes to trust anything in that world, what each one means for an AI agent, and why an agent cannot supply them by itself.

Jamey Canterbury · Co-founder and CEO, Pyrana · October 5, 2026 · 8 min read

In twenty years of audit and risk work, most of it inside life sciences companies, I never once asked a client whether they trusted their ERP system. It would have been a strange question. Nobody trusts an ERP. (People have feelings about their ERP, but trust is rarely one of them.) What we asked was whether the system was validated, who had access to what, whether the audit trail was switched on, and what happened the last time something went wrong.

So it has been a little odd to spend the last two years listening to smart people ask "Can we trust AI?" as if trust were a mood.

Trust is a deliverable

In a regulated environment, trust is something you build, document, and hand to an inspector. You do not feel your way into it.

FDA's rule on electronic records has said so since 1997. Section 11.10 of 21 CFR Part 11 lists the controls a system needs before its records count for anything: validation, access limited to authorized individuals, authority checks, and a secure, time-stamped audit trail. None of that depends on whether anyone likes the system.

I have been fortunate to contribute to ISPE's GAMP guidance over the years, and the habit GAMP drills into you is to ask what a system is for before you ask whether it works. Back in 2020, Sion Wyn and I wrote a piece for Pharmaceutical Engineering on applying that thinking to robotic process automation. The bots were a lot dumber then. The questions were the same.

They are still the same now that the bot can read a contract and argue with you about it.

Eight things, and none of them are new

Strip away the vocabulary of any one regulation and you are left with eight things you need to know before you rely on an actor, a process, a system, or a decision. I did not invent these. Anyone who has sat through an inspection could write the left two columns from memory. The right column is what each one turns into when the worker is an agent.

RequirementThe questionWhat it means for an agent
PurposeWhat is it intended to do?A defined context of use and a bounded workflow
AuthorityIs it allowed to do it?Role-based permissions, approved tools, policy gates
CompetenceCan it do it reliably?Validation, benchmarked performance, task-specific testing
ControlIs it operating within defined boundaries?Guardrails, deterministic workflow constraints, exception handling
TransparencyCan we understand what happened?Execution logs, a step-by-step trace, source attribution
TraceabilityCan we reconstruct the evidence trail?Versioned prompts, models, data, tools, context, and decisions
AccountabilityWho is responsible for the outcome?Human approval, named ownership, audit-ready records
OversightHow are errors detected, reviewed, and corrected?Monitoring, error detection, CAPA and change control

FDA is already pointing the same way for AI. Its January 2025 draft guidance on using AI to support regulatory decisions is built around a defined context of use and a risk-based view of credibility, and its 2025 guidance on computer software assurance takes a risk-based approach to the software a quality system depends on. I read both as the agency saying what it has always said. Show me the control. Show me the record.

An agent cannot hold these. An application can.

Here is where most agent projects go sideways. Teams take that list and try to make the agent satisfy it, usually by writing a longer prompt.

Take the first, second, and seventh rows. An agent on its own has no purpose beyond whatever the last prompt asked for. Its authority is whatever tools and credentials somebody handed it, and its limits are instructions it may or may not follow. Its accountability is whoever happened to be typing, with no record of what was asked, what was done, or who checked.

An agentic application is a different animal. It is a governed process with an agent inside it. People and agents work on the same task under the same roles, the same permissions, and the same record. The application defines the job. It checks every operation against the caller's role before the operation runs. It knows who requested the work, who reviewed it, and who owns the outcome. The agent does part of the work, and it is a participant like any other.

That is the whole reason an application is easier to trust than an agent. Purpose, authority, and accountability are built into the process, where you can inspect them and test them. They are not resting on how a model happens to behave on a Tuesday.

Put the controls where the model cannot argue with them

Control and transparency work the same way. In our platform they live in four layers, each with one job.

Orchestration runs the process as a fixed sequence of steps and decides when the agent is called and when a person is. The harness runs the model for one step, supplies its instructions and context, caps what it may do, and handles failures. A skill is a method for a kind of work, and it grants no access. A tool is the only way the agent touches a system, and every call is checked before it runs.

Each of those layers writes to the log. None of them can be talked out of its job by a clever prompt, which is more than I can say for most of us. My colleague Sam Merkovitz covers the harness in detail in Enterprise AI? It's all about the harness.

Competence comes from what you give it

A model does not show up knowing your procedures, any more than a new hire does. I have written before that you should not train your model to be an expert. You give it the knowledge and you manage that knowledge the way a quality system manages controlled documents: versioned, owned, access-controlled, and cited at the exact clause.

We have a small demonstration of why this matters, built on a synthetic inspection report. The task is to propose redactions under a five-rule procedure. One sentence mentions a lab technician named Amy Hamilton, and a few words later, a Hamilton syringe pump. In the first run the agent redacts "Amy" and leaves "Hamilton" sitting there, because it read the surname as the equipment brand. In the second run, same model and same document, the rule about names that look like equipment is supplied as a versioned Context Unit. This time the agent proposes the full name, flags it at 0.62 confidence, sends it to a person, and cites the rule it relied on.

The model did not get smarter between runs. It got better instructions, and it left a trail. That is competence and traceability in one example, and neither came from the model.

Competence still has to be shown, of course. You test the agent on representative cases before you rely on it, and again whenever the knowledge or the model changes.

Oversight is CAPA with better record keeping

The eighth requirement is the one that keeps the other seven true over time, and for agentic work it has a familiar shape. A reviewer finds a miss. The finding becomes a proposed change to the shared knowledge. The owner of that knowledge approves it. A new version is issued, and the original run keeps citing the version it actually used. The next run, by a person or an agent, starts from the corrected basis.

If that sounds like corrective and preventive action with change control, it should. It is the same loop, and the agent's work improves the same way a regulated process does. I made the longer argument for knowledge that gets stronger from being wrong earlier this year.

Where we are, in plain terms

Pyrana is built to make these eight things properties of the platform. Identity and role-based authorization are checked on every operation. Approval policies hold consequential actions for a named person. Context Units are versioned, owned, and cited. Every run leaves a record of who acted, under what authority, on which inputs, through which steps.

I would be a poor auditor if I stopped there. The part we are still building is the evidence for competence over time: systematic evaluation across representative cases and monitoring for drift. Per-run checks and flagged uncertainty exist today. The longitudinal picture is in development, and I would rather say that than have you find out in a diligence call.

Try it on one agent this week

Pick an agent your organization already runs and walk the eight rows.

  1. What is it for, and where is that written down?
  2. What is it allowed to do, and what enforces that?
  3. How do you know it does the job reliably?
  4. What stops it when it goes outside its limits?
  5. Can you see what it did, step by step?
  6. Can you say which version of which rule, data, and model produced a given result?
  7. Who answers for the outcome?
  8. When it gets something wrong, what changes, and who approves the change?

If you are stuck by the third question, the missing piece is the application around the agent.

Asking whether we can trust AI gets us an argument. Asking whether we can govern agentic work to the same standard we already govern regulated work gets us a to-do list. I know which meeting I would rather be in.

Sources

More for ai and data leaders

Bring your hardest process.

We will run it through admission, approval, and the record, and show you what comes out.