Ask a chatbot a question about your own procedures and it will give you a confident, fluent answer. That is the easy part, and it is also the trap. If you run regulatory affairs or quality in a regulated company, the fluent answer is not what you are accountable for. You are accountable for whether you could sit across from an inspector, point to that answer, and show how it was produced and who signed off on it.
That is a higher bar than "the answer sounds right," and most AI tools were never built to clear it. This is about what separates an AI answer you can defend from one you cannot, using the standard you already live by rather than a new one.
The only test that matters is the audit
In a regulated setting, the value of an answer is capped by whether it holds up under review. A brilliant answer you cannot trace, log, or attribute is worth less than a plain one you can, because the plain one survives an inspection and the brilliant one becomes a finding.
So the useful question to ask of any AI system is not "how good are its answers." It is "if this answer ended up in a decision an inspector later questioned, what could I show them." If the honest reply is "the model said so," you do not have a tool you can use on regulated work. You have a liability with a nice interface.
A plausible answer is not a defensible one
Language models are built to produce fluent, plausible text. That is exactly the property that makes them useful and the one that makes them dangerous in regulated work. A model can state something that reads perfectly and is wrong -- a behavior usually called hallucination. It can also be right for reasons it cannot show you.
Neither case survives an audit. An inspector does not grade prose. They ask where a claim came from, who could change it, and how you know it is current. A plausible answer with no provenance fails all three questions at once. The goal is not a system that sounds more convincing. It is a system whose answers carry the evidence of how they were produced.
Borrow the standard you already live by
You do not need a new framework to judge AI output. You already hold your records to one, whether you build medical devices or make pharmaceuticals. Two pieces of it map directly onto what a trustworthy AI answer requires.
Under FDA 21 CFR Part 11, electronic records that support regulated decisions need secure, computer-generated, time-stamped audit trails that capture who did what and when, retained as long as the records themselves and available for the agency to review. Access is limited to authorized individuals. The FDA's 2018 data-integrity guidance frames the same expectation through ALCOA -- the shorthand for data that is Attributable, Legible, Contemporaneous, Original, and Accurate, later extended to add Complete, Consistent, Enduring, and Available.
Read those requirements again with an AI system in mind. Attributable. Traceable. Access-controlled. Reviewable. The standard for a defensible answer is not exotic. It is the standard you already apply to every other record that touches an inspection. None of this makes an AI system "Part 11 validated" on its own, and it is not legal advice. The point is narrower. The properties that make a record defensible are the same ones that make an AI answer defensible.
The standard for a defensible AI answer is not exotic. It is the standard you already apply to every other record that touches an inspection.
The four properties a defensible AI answer needs
Strip it down and four properties do the work. A system that has all four produces answers you can stand behind. A system missing any one of them puts the accountability back on you with nothing to show.
1. It cites its source
Every answer should point to the specific passage in your own documents that it came from, so a person can open that source and confirm the answer against it. This is what makes an answer attributable and verifiable instead of a claim you have to take on faith. A cited answer turns the reviewer from a believer into a checker, which is the entire point. If a system cannot show its source, treat its output as a rumor, not a record.
2. It is logged
The system should keep a secure, time-stamped record of who asked what and what it returned -- the same shape of audit trail your other regulated systems already keep. That log is what lets you reconstruct, months later, how a given answer was produced and used. Without it, every answer is an orphan with no history, and "how did we arrive at this" has no answer.
3. It is access-controlled
Not everyone should see everything. The system should enforce role-based access control, meaning it only shows a given user the material they are already cleared to see, rather than flattening your permission structure the moment documents are indexed. This is both a data-integrity property and a security one, and it is one of the first things a careful reviewer will probe.
4. A human reviews and approves
The system drafts and retrieves. A qualified person reviews, corrects, and approves. Accountability stays with the person, where the regulation already puts it, and the AI is a fast first pass rather than an unaccountable decision-maker. Human-in-the-loop is not a hedge. It is the design that keeps the output defensible, because a named person can always say what they checked and why they accepted it.
Where cloud chatbots fail this test
Hold a general cloud chatbot up to those four properties and the gaps are immediate. It rarely cites the specific source passage behind an answer, because it was built to respond, not to attribute. Its logs, if any, live on the provider's side and are not yours to inspect or hand to an auditor. It does not know your access model, so it cannot enforce it. And its default posture is to answer, not to route the answer through a named reviewer.
That is before the separate problem that sending your regulated documents to a third-party cloud may not be permitted at all. Even setting the data-residency question aside, the ordinary chatbot fails the audit test on its own terms. It was designed for a different job.
What to check before you trust an answer in regulated work
If you are evaluating any AI system for regulated work, these questions separate a defensible tool from a demo. They are worth putting in front of whoever owns data integrity on your side.
- Does every answer cite the specific source passage, so a reviewer can verify it against our own document?
- Is there a secure, time-stamped log of who asked what and what the system returned, and is that log ours to keep and inspect?
- Does the system enforce our access controls, so it never shows a user material they are not cleared to see?
- Does a qualified person review and approve before an answer is relied on, and is that step recorded?
- Can we reconstruct, after the fact, how any given answer was produced?
A system that answers those cleanly gives your team a faster way to work that still holds up under review. A system that cannot is asking you to put your name behind output you cannot defend. In regulated work that is not a trade worth making, no matter how good the answer sounds.
Judge an AI answer the way you judge a record, not the way you judge a search result. Held to that standard, the tools that survive are the ones built to show their work.
Metellus Partners builds private, on-premise AI for regulated industries, with cited answers, audit logging, and access control built in, running inside the client's own infrastructure. If defensibility is the bar you have to clear, get in touch.
Sources
- U.S. FDA, 21 CFR Part 11 (Electronic Records; Electronic Signatures) -- on secure time-stamped audit trails and access limits for electronic GxP records.
- U.S. FDA, Data Integrity and Compliance With Drug CGMP: Questions and Answers (2018) -- on ALCOA data-integrity expectations.