Who Evaluates the Evaluators?
AI Assurance Needs Infrastructure. Here’s How To Build It.
In our last piece, The Missing Layer, we talked with enterprise AI deployers across sectors about what it’s like to manage AI systems in the absence of shared trust infrastructure. The picture was consistent: sophisticated organizations doing impressive work, largely from scratch, without the shared standards, independent verification, or institutional backing that would let them demonstrate trustworthiness in ways the market can recognize.
This piece is about the other side of the equation: the people and organizations whose job is to independently evaluate whether AI systems are safe and trustworthy to use across a variety of use cases. The field is growing, but it’s still young and it doesn’t amount to the trust infrastructure AI demands.
I. What Is AI Assurance, and Why Does It Matter Now?
AI assurance is the process of independently measuring, evaluating, and communicating the trustworthiness of AI systems and of their impacts. It produces evidence that deployers, regulators, insurers, and the public can rely on. The key word here is independent. This is not the vendor checking its own work.
In practice the work varies enormously.
Some providers specialize in TEVV (Testing, Evaluation, Verification and Validation): the umbrella term for technical assurance work that consists in systematically testing how a system performs—from red-teaming it to find where it breaks, or benchmarking it against a standard.
Some deploy domain experts, such as doctors, hiring managers, or insurance adjusters, to evaluate whether a system would perform reliably in their field.
Others build platforms that let organizations run and manage their own testing.
Many do a combination of all three and more. And often, the system being evaluated is one no standardized test was ever designed for. Dr. Rumman Chowdhury, founder and CEO of the evaluation infrastructure startup Humane Intelligence PBC, offers two examples from her own practice:
“This could be an insurance product, and a provider that is using an AI agent to determine whether or not someone’s contesting of their insurance is valid. There is no specific, clean and easy benchmark that is testing that agent. Another example might be an AI being used in a car to help drivers with navigation. Again, there are no specific out-of-the-box tools that can solve that problem, so we help orchestrate that environment.”
What unites the field is a shared purpose: producing credible, independent evidence about how AI systems actually behave, so that the organizations deploying them, the regulators, and the public can rely on them instead of on a vendor’s self-assessment.
The demand for this work is arriving from several directions at once. Enterprise AI spending reached $37 billion in 2025 and is projected to accelerate. The deployers we spoke with in our previous piece were clear: they need independent assurance they can point to, and they don’t yet have a reliable way to find it or judge its quality.
II. What Makes AI Different
The deployer piece laid out three properties that make AI systems harder to govern than conventional software: non-determinism, opacity, and autonomy. Those same properties define the technical challenge for the people doing the evaluating.
Non-determinism means you can’t run a test once and declare the system verified. The same input can produce different outputs depending on the model version, its settings, or plain chance. Evaluation has to be continuous and repeated, and it has to keep pace with updates. A system that passed evaluation last month may behave differently today.
Opacity means evaluators often work with limited visibility into how a model arrives at its answers. Traditional software auditing assumes you can inspect the code. With many AI systems you can’t, and how much access an evaluator is granted ends up shaping the entire methodology.
Autonomy means the system’s behavior unfolds over chains of decisions rather than single answers. An AI agent that books travel or processes claims can’t be judged one output at a time. You have to evaluate the whole sequence, including what happens when something unexpected occurs partway through. The thing being evaluated is a process and an output, not a product.
In this environment, evaluators are frequently building the test as well as administering it. Chowdhury describes an automotive client worried about a risk no benchmark covers:
“If, in this driving example, they’re concerned about this AI system leading to distracted driving, we then have to come up with a test that enables them to test distracted driving. We’re translating words into code, basically.”
Agents pose the hardest version of this problem. Dr. Shea Brown, founder and CEO of the AI assurance firm BABL AI, points to the sheer number of situations an agent can be exposed to, and the absence of any settled way to test them:
“Agents are the hardest things to evaluate and to test, because the parameter space is so big. … At the end of the day there’s some output which might have some consequential impact. How do you test that? It’s very, very difficult to do a repeatable test. If the companies themselves haven’t set up that infrastructure, then as an external auditor you are at best doing red teaming, or you have to test little individual components at a time, and then in aggregate you have to say: okay, given the sum, we think it’s probably safe to use. And that’s just really difficult.”
And then there is context. The same foundation model can sit behind customer service, clinical recommendations, financial underwriting, or content moderation. Each use case carries its own risks, its own failure modes, its own definition of “good”. Evaluation has to be specific to the deployment, which means domain expertise matters as much as technical skill. Chowdhury again:
“It could be the same product, and I’m marketing it to an under-18 category versus people who are adults, and that will completely change the game. Or maybe geographically, or people with disabilities. It just looks totally different. … It’s not just about being an AI engineer or an expert in AI. It’s actually a lot of domain expertise that’s needed.”
These properties don’t just make AI harder to evaluate. They make AI evaluation a new discipline: one that borrows from software testing, statistics, and traditional audit and assurance disciplines but cannot be reduced to any of them.
III. AI Assurance in Historical Context
Every recognized assurance profession was built over time through hard-won experience, not designed in a paper specification. The infrastructure that makes independent evaluation credible in other fields was constructed over decades, usually in response to market failures and crises.
Financial auditing is perhaps the closest analogy. Before the Securities Acts of 1933 and 1934, companies made claims about their financial health and investors had no standardized way to verify them. Generally Accepted Accounting Principles (GAAP), oversight from the Public Company Accounting Oversight Board, CPA licensing, and eventually Sarbanes-Oxley, passed after Enron and WorldCom, built the machinery capital markets now take for granted: shared standards, certification, codes of conduct, clear liability.
Product safety tells a similar story. Organizations like Underwriters Laboratories (UL), Bureau Veritas, and SGS created the testing and certification infrastructure that consumers rely on without a second thought. The UL mark on a toaster is precisely the kind of trust signal AI systems lack: a recognized, independent indicator that a product has been evaluated to be safe. The global testing, inspection, and certification industry is worth over $400 billion—proof of the scale of market that trust infrastructure can create.
The analogies only take us so far. A financial auditor examines a fixed set of books at a defined point in time. A product safety tester evaluates a physical object that generally behaves the same way on Tuesday as it did on Monday. AI assurance has to evaluate a probabilistic, continuously updated system deployed in varying contexts. The institutional skeleton of the older professions translates well: standards bodies, accreditation, independence rules, liability frameworks. But technical muscles largely do not.
As Brown put it in his own writing:
“The harder challenge is cultural and institutional. It requires the AI policy community to engage seriously with a professional literature that most technologists have not read, and it requires the assurance profession to engage seriously with technical AI evaluation in ways that go beyond process auditing. Neither community has fully made that move yet. But the foundation for what we need is already there.”
The expertise the field needs is split across two different communities. One understands how assurance works: independence, evidence, professional skepticism, the discipline of standing behind a written opinion. The other understands how AI systems fail and how to catch them. Each provides only a partial solution without the other.
IV. The Institutional Infrastructure the Field Needs
The technical challenges are real, and the field’s ability to meet them depends on institutional scaffolding that doesn’t yet exist. But most can be adapted from the older assurance professions. Four categories matter most.
1. Codes of Professional Conduct and Independence
Independence is the core value proposition of third-party assurance. It is what separates a credible evaluation from a paid endorsement. In financial auditing, independence is governed by detailed rules—you cannot audit a company you have a financial stake in—and violations carry professional consequences.
Patrick Sullivan of the cybersecurity and compliance firm A-LIGN points out that independence cannot simply be asserted. It has to be validated by someone outside the relationship:
“There will always be the perception of loss of impartiality when it’s a pay-for-play business. … The third party is necessary to ensure that the first party that’s creating and placing goods on the market is evaluated consistently and fairly by the second party—audit firms, certification bodies, conformity assessment bodies, whatever we want to call them. The accreditors sit above that and ensure that everyone is playing the game by the right rules.”
AI assurance has no equivalent rules yet. Many evaluators also hold commercial relationships with the very companies whose systems they might evaluate. The field lacks shared definitions of independence, conflict of interest, and transparency. Without them, every evaluation’s credibility is ad-hoc, and rests on the reputation of the individual or firm rather than on structural guarantees that make trust scalable. Chowdhury is blunt about the consequence:
“Being a purely independent evaluator is nearly impossible, because it’s just not a fully built-out ecosystem yet. But it needs to be.”
2. Standards of Practice
How should assurance work be conducted? What counts as a rigorous evaluation? Right now, every provider answers differently. As Chowdhury noted in The Missing Layer, a benchmark is a test, not an evaluation: a real evaluation defines scope, chooses methods, analyzes results, and reports them. None of that is standardized across the field.
Sullivan says the confusion starts even earlier, with what is being evaluated in the first place:
“We can say ‘AI assurance,’ ‘AI audit’—that’s so semantically charged it’s almost impossible to know what we’re actually talking about. Are we talking about a governance audit? That’s one set of activities. Are we talking about evaluating for disparate outcomes, for some unintended bias? That’s a fundamentally different thing. … A clean description of the target of evaluation is noticeably absent from what we do today.”
Fixing this doesn’t mean mandating specific tools or a single method. It means setting shared expectations for what a complete evaluation includes, so the deployers commissioning the work and the regulators relying on it know what they’re getting. As Chowdhury puts it:
“It’s not necessarily about being prescriptive—run this exact test—but to say: your assurance evaluation should accomplish these specific goals.”
3. Accreditation and Certification
Two questions follow. At the organizational level: how does a deployer know an assurance provider is qualified? At the individual level: what credentials should an AI assurance professional hold? As Chowdhury put it, “in order to have individuals who are considered experts, we do have to have something to assess them against.”
Brown names accreditation as the piece of infrastructure he would build first:
“It’s hard for clients—I feel bad for clients. They can’t really tell who’s legitimate and who can actually do the work. … There’s no way of telling externally whether a company is equipped to do this kind of work. You can look at people’s LinkedIn profiles and they sound amazing, but have they ever really done any of this?”
Today the field runs on reputation and personal networks—a Fortune 100 company looking for AI evaluation calls one of the handful of recognized names. That doesn’t scale, and the credentials that do exist often prove very little. Accreditation at both the organizational and individual level would give deployers a reliable way to find qualified evaluators, give evaluators a way to demonstrate competence, and give the profession a way to hold the line on quality as it grows.
4. Professional Liability Clarity
Who is liable when an AI system that passes an independent evaluation later causes harm: the evaluator, the deployer, the developer? These questions are almost entirely unresolved. Until they are, evaluators face open-ended liability that discourages entry, deployers can’t lean on evaluations as part of a legal defense, and insurers have little basis for underwriting the work. Brown puts specialized insurance near the top of his list:
“Having good specialty insurance for auditors and assurance providers in this space would be good: standard things that we can trust and know are going to work, as opposed to just general E&O or liability insurance. Something that’s really built for the profession.”
Chowdhury has named legal protection for independent evaluators as one of the most useful things legislation could provide. The point isn’t to shield anyone from the consequences of shoddy work. It’s to create the clarity that lets a market function: defined obligations, liability, and recourse.
V. The Technical Infrastructure the Field Needs
Institutional scaffolding enables the profession; technical infrastructure enables the work itself. Here, the newer half of the field—the researchers and engineers who study how AI systems fail—carries most of the load. Four layers are needed, each building on the one before.
1. Evaluation and Measurement Science
Underneath everything sits a scientific question: how do you measure a system that’s non-deterministic, opaque, and context-dependent? In many areas, the honest answer is that nobody yet knows. Important work is underway—the NIST AI Risk Management Framework, academic research on evaluation methods, industry initiatives—but the science is young.
That’s worth saying plainly. Some of what matters most to deployers and the public—long-term societal effects, emergent behavior in agentic systems, subtle patterns of unfairness—are among the hardest to measure. A field that admits the limits of its current tools is more credible, not less. Chowdhury’s own list of open problems starts with time horizons:
“I think a lot about evaluation of short-term products versus long-term impacts. Social media is a perfect example of optimizing for short-term outcomes that lead to long-term harms. … Now we are talking about things that are very, very real: cognitive offloading, critical thinking, what we mean by workforce augmentation. These are not abstract questions. These are fundamental questions of AI applications that are not easily picked up in pushing a product out the door.”
2. Standards: Technical Criteria for Evaluation
In financial auditing, you test against GAAP. In product safety, you test against a published standard. In AI assurance, the criteria today are usually worked out with the client, engagement by engagement: what does safe mean for this system, in this context, for these people? Chowdhury describes how consistently that falls to her:
“A hundred percent of the time, I am co-working to help define good. Right now we do not have very clear standards. … The big issue here is being able to scale. The comforting thing about standards and norms is that people can say, okay, just run with the standard — and then we can start building tooling and standardized testing around these kinds of things.”
That co-creation produces meaningful, context-specific evaluation. It also does not scale. Every engagement starts from scratch, and no two firms’ results can be compared. As Brown puts it:
“It’s still the Wild West. There are no hard standards that say: this is what sufficient testing looks like, or this is what a sufficient safety metric should be.”
The field needs baseline technical criteria — at minimum, working definitions of safe, fair, reliable, and transparent for AI deployed in healthcare, financial services, consumer products, and other high-stakes settings. Baselines should be floors, not ceilings: the minimum an evaluation must address, with room left for context-specific depth.
3. Methodologies: How to Apply Evaluation Against Standards
Red teaming, automated testing, human-in-the-loop review, benchmark suites, stress testing, continuous monitoring—these are the tools in the evaluator’s kit. What the field lacks is a shared framework for when to use which, how to combine them, and how to judge the quality of the methodology itself.
In practice, a consensus is starting to form on its own: sophisticated clients combine several methods rather than trusting any single test. Chowdhury describes the pattern:
“The trend I’m seeing in industry is what I call stacking their evaluations. They’re not evaluating a product on one test—they’re running a series of tests. … For me, an evaluation starts with a real design: here’s what good is, here’s what I want to measure, here’s what I want the outcome to be, and who I’m testing it with. Then you pick the right tool: a benchmark, human-in-the-loop red-teaming, LLM-as-judge.”
This reflects an emerging consensus that robust evaluation requires multiple lenses. Codifying it into shared methodological frameworks would help the field move from craft to profession.
4. Tools and Shared Assets
Finally, supporting all of the above: the software libraries, datasets, benchmarks, and platform providers need to do the work efficiently and consistently. Open-source frameworks like HELM and Inspect are a real start. What’s missing is the connective tissue: benchmark suites calibrated to specific domains, common formats for reporting results, platforms that work together rather than in isolation.
Brown’s concern is comparability. Without shared reference points, every firm’s numbers stand alone:
“You really want to be following a standard if you want to provide assurance. There either need to be some benchmarks, or there needs to be a very scoped way of doing that kind of testing. If everybody comes up with their own thing, it’s really hard to compare numbers across firms and keep independence.”
There is also a simpler problem: buyers often can’t find what already exists. Part of Humane Intelligence’s work, Chowdhury says, is matchmaking:
“Imagine somebody in a Fortune 100 company whose day job isn’t AI evaluation. It’s hard to know exactly what to do, given how fast the industry moves. We help improve discoverability, and match people to the tools that solve their problem.”
The analogy here is clinical trial infrastructure: the registries, protocols, and data standards that let pharmaceutical results be compared and built on across institutions. AI evaluation needs the equivalent, so bespoke engagements can become cumulative practice.
VI. The Path Forward
The infrastructure described here—institutional and technical—is too large for any single organization to build. It requires coordination among providers, deployers, insurers, academics, and civil society. Brown’s read on the moment is that the pieces are close:
“Everything that’s needed is in germination right now. The standards are germinating, the professional organizations are germinating, and the legal impetus — the risk impetus — is going to start generating demand pretty soon. It’s just making sure that there are companies and people who are ready to go when the dominoes start falling.”
Some of that coordination is already starting to take shape, quietly, among the providers, deployers, and researchers we’ve been in conversation with. We’ll share more on what that looks like soon.
AI assurance has the chance to build proactively, ahead of the failure that would otherwise force the issue. The deployers are asking for it. Regulators are beginning to require it. The technology demands it. The open question is whether the assurance community organizes to build it together—or leaves it to be defined by default, by the largest incumbents, the most aggressive regulators, or the next headline failure.

