The AI Auditor: the role production AI has been missing | Scorable

Updated: 2026-05-15
By: Ari Heljakka

The tedious problem

Here is what operating production AI in a regulated environment looks like right now, in most organizations:

Meanwhile, the automated test suite passes. Latency is fine. Error rates are within SLA. The system looks healthy by every operational metric, and nobody has a credible answer to the question: "Is it actually doing what we said it would do?"

This is not just a tooling problem. Better evaluation tooling helps, but it does not close the gap, because the team that built better tooling is the same team that chose which behaviors to test, which thresholds to set, and what counts as a passing score. A tighter version of that process still inherits the same incentives. The gap is not that the engineering team is careless or dishonest. It is that the engineering team cannot independently verify their own choices. That is a structural problem, and tooling cannot solve a structural problem.

What production AI is missing

Production systems have had to invent new functions before when old structures stopped working; that instinct is the right one here. The analogy we will lean on is older and tighter: financial audit.

A company's accounting team is skilled, well-intentioned, and deeply familiar with the numbers. They still cannot audit themselves. The reason is not moral; it is procedural. The team that produces a set of accounts has an inherent conflict of interest in evaluating whether those accounts are accurate. The audit function exists to close that conflict by design, not by goodwill.

The same logic applies to AI systems. The engineering team that selected the model, designed the prompt scaffolding, chose the evaluation thresholds, and built the test suite cannot independently verify that those choices produce trustworthy outputs. The problem here is not primarily one of intent; it is one of orientation. Engineers building an AI system face two structural blind spots:

The gap is not competence; it is structure. And the structure requires a function that sits outside the engineering organization, watches what the system actually does in production, and produces a record the people who built the system cannot quietly edit.

Call this function the AI Auditor.

What an AI Auditor does

The AI Auditor function has four components. Together they form something that does not exist in most organizations today: continuous, independent, evidence-grade evaluation of production AI.

What it is not

The AI Auditor function is distinct from several things organizations may already have:

What changes when the role exists

The absence of an AI Auditor function produces a particular kind of organizational dysfunction: accountability theater. The engineering team knows the system has gaps; they fix what they can find. Compliance asks whether the system is safe; engineering says yes, with caveats. Nobody is lying, but nobody is producing the kind of structured, independent record that would let a third party reach an informed conclusion.

When the function exists, several things change:

The regulatory moment

The external pressure is arriving at the same time as the internal need. The regulatory frameworks differ in emphasis, but they share a common requirement: records that someone other than the team that built the system can use to reach an independent conclusion.

Consumer Duty wants outcomes evidence. Firms must show that products and services are actually delivering good outcomes for customers, not just that the product was designed with good intentions. An AI system that advises, recommends, or decides on behalf of customers needs continuous documentation of what it actually did, not a retrospective claim.

The EU AI Act, for high-risk applications, sets out requirements for technical logging and traceability (Article 12), alongside post-market monitoring and record-keeping obligations, and oversight measures where explicitly required. Records that exist only on paper, or that the development team alone controls, do not meet the intent of the requirement.

MiFID II treats AI-assisted investment advice the same as human advice for record-keeping purposes. The evidence trail has to be equivalent, which means it has to be continuous, versioned, and available for examination.

The common thread is not that these frameworks all require a third-party auditor. The common thread is that they all require records a third party can actually use. The AI Auditor function is what makes that possible.

Where we go from here

The AI Auditor is not a role that replaces what engineering teams do. It is a role that makes engineering-team evidence credible to people outside the engineering team. That distinction matters because most of the pressure arriving on AI-deploying organizations right now is coming from people outside the engineering team: boards, regulators, risk committees, procurement reviewers.

Building this function is not simple, and it is not cheap. It requires instrumentation, access control, change-management process, and the discipline to maintain a rubric even when it surfaces uncomfortable findings. What it produces, when it works, is something genuinely new in most organizations: the ability to say with evidence, not just assertion, that your AI system is doing what you said it would do.