AI Agent Certification: What AIUC-1 Actually Proves
Do I need an AI agent audit? If you buy or deploy agents that touch your data, ask for evidence, not a badge. AIUC-1 — the standard its own company markets as the “SOC 2 for AI agents” — is a third-party report on how one agent behaved against a defined scenario set at a point in time. Worth reading; not proof your deployment is safe.
What is AI agent certification?
An outside firm examines an agent against a written standard and reports what it found, so the vendor cannot grade its own homework. The limit is scope: the firm only sees what was put in it. AIUC-1, launched on 17 July 2025, is a security, safety and reliability standard across six areas — data and privacy, security, safety, reliability, accountability and society — requiring more than 50 safeguards plus third-party testing, and shaped by more than 250 security and risk leaders at Fortune 1000 companies per the company’s current materials (AIUC; the standard’s site).
The AIUC-1 standard: what it actually tests
The tests are adversarial, not documentary. AIUC says certification runs an agent through risk-and-attack scenarios in three classes. A jailbreak makes an agent ignore its own rules; the dangerous version is indirect, hidden in content it was asked to read. A hallucination is confident output that is not true — worse for an agent than a wrong answer, because the wrong answer can be acted on. A data leak is information leaving through a prompt, a tool call or an over-broad permission.
Scale is a range: the funding release cites 5,000 combinations tailored to the business type, the standard’s own site says red-teaming “typically involves 1,000 to 5,000 test scenarios”, and published audits ran 900+ (KPMG) and, from a single outlet, 2,000+ (UiPath). The tests are buyer-defined: AIUC asks its consortium, not vendors.
Is AIUC-1 the “SOC 2 for AI agents”?
That is the company’s framing: marketing language, not technical equivalence. AIUC’s table puts an “audit report with certificate” for AIUC-1 against an “attestation report” for SOC 2, both displayed for 12 months, and contrasts AIUC-1’s forward-looking adversarial tests with SOC 2’s backward-looking assessment. AIUC also says AIUC-1 “does not duplicate the work of non-AI frameworks like SOC 2, ISO 27001, or GDPR”. Not regulator-endorsed; no law requires it.
What a 100-page audit report proves, and what it does not
TechCrunch reports that AIUC’s testing produces a roughly 100-page report on where an agent performs safely and reliably, and where it does not. Read it as evidence of tested behaviour on a defined scenario set at a point in time: not a guarantee the agent never fails, and not coverage of your configuration or data — Cursor’s certification ran against its key product surfaces on a representative configuration, not your deployment. A material finding produces a qualified or adverse report and re-testing.
How often should AI agents be audited?
Quarterly technical testing under a 12-month certificate: the term is 12 months, operational controls are reviewed annually, testing and retesting happen at least quarterly, and the standard is revised quarterly — 1 January, 1 April, 1 July, 1 October. The funding release compresses that to “independently audited and recertified every quarter”; the annual renewal is what keeps the mark valid. A certificate issued last year describes last year’s configuration. This site treats an audit as a loop: see the step-by-step agent security assessment and the AI agent risk checklist.
Who audits the auditor’s automation?
The recursion problem is stated, not hidden: AIUC uses AI agents to run the tests and AI to analyse the data, with humans verifying the audit. That is the same question this site raises about agents gaming their own benchmark scores — any automated evaluator becomes something to be optimised against — and about agents running code nobody wrote, where the toolchain performing the check is itself attack surface.
Who is certified today, and what to do if your vendor is not
TechCrunch names four: Cursor, Lovable, Harvey and ElevenLabs. KPMG LLP, UiPath and Fin (formerly Intercom) appear in the company’s own materials; KPMG’s covers one platform, aiQ Capture, the first Big Four capability certified. Schellman became the first authorised AIUC-1 auditor on 3 February 2026. If your vendor is not listed, ask for the report anyway, then scope, date and re-test schedule — not a logo. ElevenLabs certified at launch and used it for what AIUC calls first-of-its-kind agent insurance.
AI agent insurance vs certification: three different things
Certification, insurance and a warranty transfer different things: a certificate is evidence, a policy is financial protection, a warranty is a promise enforceable between the parties.
| AIUC-1 certification | Self-attestation | Insurance | |
|---|---|---|---|
| What it is | A report plus a 12-month certificate | The vendor’s own statement, its own criteria | A policy that responds after a loss |
| What it evidences | Tested behaviour on a defined scenario set, at a point in time | That the vendor says it did the work | Nothing about how the agent behaved |
| What it does not cover | Your configuration and data | Anything the vendor chose not to test | Whether a loss is a covered claim |
Coverage is separate: see what current AI agent policies do and do not cover. If an agent of yours touches a third party’s system, the RubyGems agent attack shows why scope and log retention get asked first.
What to ask your agent vendor
When a vendor claims certification, ask which standard, who ran the tests, what was in scope, when it was last run and when it is next due:
- Who ran the tests? Schellman was the first authorised AIUC-1 auditor; the standard’s site also names Coalfire.
- What was in scope? Which product, surfaces and configuration — KPMG’s covers one platform, not the firm.
- Will you share it? A certificate without the report is a logo, not a finding.
Agency buyers: what to demand in the contract. For the difference between a certification the vendor buys and an evaluation the vendor funds, work through the buyer’s test below.
Sept 2026: embedded vs independent evaluation
Two different things get sold as an AI evaluation, and the difference is who pays. An embedded evaluation is funded and staffed by the vendor: the company being judged pays an evaluator to work inside it, with access comparable to an employee. An independent evaluation is commissioned by a party with no financial tie to the vendor — someone the vendor does not pay, cannot stop, and does not approve the findings of. Both can produce real findings. Only one answers the buyer’s question: can this evaluation say no to the company it is judging?
Who evaluates the evaluator here?
For the arrangement Anthropic announced on 18 September 2026, Anthropic pays Accenture’s Faculty to evaluate Anthropic, and no party to it is named as evaluating Faculty. Anthropic’s own post says there are, as yet, no standards for what an embedded evaluator may access or how it should report what it finds, so nothing published gives a second evaluator anything to check the first one’s work against. That is the answer to the question in the heading; the rest of this section is the evidence for it.
The partnership is with Accenture, on what Anthropic calls independent evaluation of frontier AI, as a step toward the commitment it made to embed evaluators inside the company. The work is led by Faculty, described in the announcement as Accenture’s specialist AI business, and Anthropic names the scope: “evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards”. Faculty is the AI division of Accenture, the consultancy that acquired it in January to act as that division — TechCrunch, which calls Accenture a technology consulting giant, gives no year for the acquisition and dates the article 18 September 2026. TechCrunch’s headline calls Faculty Anthropic’s “first embedded evaluator”; that ordering claim is TechCrunch’s, because Anthropic says the partnership is non-exclusive and that it is in dialogue with METR and other nonprofit evaluators piloting embedded evaluation using their own funding.
Who pays is stated by Anthropic, in one sentence. “Given the importance and urgency of this work, Anthropic will fund Accenture’s work directly.” The evaluator is paid by the company being evaluated, which is what embedded means. Anthropic’s case is that this makes its accountability verifiable rather than smaller: “To be clear, independent embedded evaluators do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility.” It also concedes what is missing: “There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find.” Read plainly, independent here means independent of the lab’s own research teams — not disinterested, and not clear of the lab’s money. It is also not an audit against a standard: Anthropic’s post names no standard, because there is none for this.
The $1 billion figure, read properly. Anthropic’s post says: “Anthropic and Accenture each expect to invest at least $1 billion in building capacity in this area over the next five years.” That is an expectation of investment, by each party, across five years — not money spent, not a signed contract, and not a published combined total. TechCrunch renders the same figure without the word each (“Both companies expect to invest at least $1 billion in the project over the next five years”), which is the one rendering that implies a single total.
The counter-position, published the same day. CNBC reported that “Over 100 artificial intelligence experts and evaluators are banding together to warn they won’t have the necessary resources and protections to test the safety of AI technology”. The letter, organised by the AI Evaluator Forum, does not oppose embedding — it opens by welcoming it: “We, the undersigned, are encouraged to see frontier AI companies call for embedding third-party organizations to evaluate rapidly escalating AI capabilities and risks.” It then sets out the four conditions it says credible embedded evaluation requires — “scientific objectivity, transparency, independence, and robust protections” — and makes the money and ownership terms explicit: an embedded evaluation organisation “should not accept any form of payment or other reward contingent on the evaluator’s findings”, should not be “owned or governed by frontier AI companies”, and should not have “other significant commercial business with them”. Signatories include Geoffrey Hinton and members of Johns Hopkins University, Stanford University and the nonprofit evaluator METR. The letter addresses frontier AI companies; the two lab names in the coverage are CNBC’s framing, not the signatories’.
Set the two side by side and the disagreement is narrower and sharper than for-and-against: the coalition welcomes embedded evaluation and asks for conditions on money, ownership and transparency, while the first arrangement it will be measured against funds its evaluator directly and, in the lab’s own words, has no standards yet for access or reporting. What a buyer can verify today is smaller than either position — who wrote the questions, and who holds the invoice.
The buyer’s test: four questions for any evaluation
A certification, an audit report and a consultancy evaluation reach your desk in the same shape: a PDF and a logo. These four questions sort them, and they work the same way on a vendor certification such as AIUC-1 and on an embedded evaluation like the Faculty arrangement above.
- Who pays the evaluator? Ask who invoices whom, and whether any part of the fee depends on what the evaluator finds. Anthropic answers this one in writing — it funds Accenture’s work directly. The evaluators’ own minimum condition is narrower than independence: an organisation “should not accept any form of payment or other reward contingent on the evaluator’s findings”. A paid evaluation is not automatically worthless; an undisclosed payment is.
- Who wrote the test? Ask for the scenario set and who defined it. AIUC’s own account is that its tests are buyer-defined — the company says it asks its consortium, not vendors. On the embedded side, Anthropic says no standards yet govern what an evaluator may see, so scope is negotiated between the funder and the evaluator rather than published. If the evaluated company writes the questions, a second signature does not turn a self-assessment into an audit: see how agents game the benchmarks they are measured on. The vendor-side version of the same question is Anthropic’s September 2026 metrics proposal — measures published so a third party could check them.
- Is the report public? A certificate without the report is a logo, not a finding. Ask for the report, its date and its scope, and for who else may read it. Nothing published says the Faculty evaluation will produce a public report, and Anthropic’s own caveat is that no standard yet governs how embedded evaluators report what they find. If you are going to be the only reader, note that in the procurement file.
- What happens on a failed result? Look for the consequence, not the promise: a qualified or adverse report and re-testing (the AIUC-1 path described on this page), disclosure to buyers, a contract remedy — or nothing at all. No source on the embedded arrangement names a deliverable, a report or a consequence of a failed result. An evaluation with no defined consequence is a subscription, not a control.
Four questions you cannot get answered are themselves an answer: if the vendor will not say who pays, who wrote the test, who can read the report, or what happens when something fails, then independently evaluated is doing more work than the evidence behind it. For putting that in writing with a supplier, see what to demand in the contract. On the agency side, what to disclose when you tell a client an agent was independently evaluated is worked through on Find AI Agency.
Questions owners are asking
Do I need an AI agent audit if my vendor is certified?
Yes, a smaller one than the vendor’s. A certificate is evidence about the vendor’s agent on a defined scenario set at a point in time; it does not cover your configuration or data. Yours: what the agent can reach, what is logged, who can stop it.
What is AI agent certification?
An outside firm examines an agent against a written standard and reports what it found, so the vendor cannot grade its own homework. AIUC-1 covers six areas — data and privacy, security, safety, reliability, accountability and society — with more than 50 safeguards plus third-party testing.
What does the AIUC-1 standard test?
Jailbreaks (inputs that make an agent ignore its rules, including indirect ones hidden in content it reads), hallucinations (confident output that is not true) and data leaks (information leaving through a prompt, a tool call or an over-broad permission). Scopes run 1,000 to 5,000 scenarios.
Is AIUC-1 the SOC 2 for AI agents?
That is the company’s framing, not an equivalence. AIUC contrasts its audit report with certificate against SOC 2’s attestation report, and quarterly testing against an annual cycle. AIUC says AIUC-1 does not duplicate SOC 2, ISO 27001 or GDPR.
How often should AI agents be audited?
Under AIUC-1: technical testing and retesting at least quarterly, controls reviewed annually, a 12-month certificate, and a standard revised on 1 January, 1 April, 1 July and 1 October. Treat an audit as a quarterly loop.
Is AI agent insurance the same as certification?
No. A certificate is evidence about tested behaviour; a policy is financial protection after a loss. AIUC says insurers are offering AI-specific coverage to certified agents, and ElevenLabs took agent insurance off its February 2026 certification.
Who audits AI models?
No single regulator does; named third parties do, and who pays them decides what the result is worth. AIUC-1 agent certification runs through authorised auditors such as Schellman (the first authorised AIUC-1 auditor, 3 February 2026), while Anthropic pays Accenture’s Faculty directly to evaluate Anthropic — the embedded-evaluator arrangement it announced on 18 September 2026.
Is Anthropic independently evaluated?
Yes — but independently of its own research teams, not of its money. Anthropic’s 18 September 2026 post says it funds Accenture’s Faculty directly to evaluate and red-team its models, and states that no standards yet govern what an embedded evaluator may access or how it should report what it finds.
Are AI safety evaluations independent?
Only when the party paying has no financial tie to the vendor and cannot suppress the findings. The AI Evaluator Forum letter signed by over 100 experts and reported by CNBC on 18 September 2026 asks for exactly that — no payment contingent on the findings, no ownership by frontier labs, no other significant commercial business with them — and a vendor-funded evaluation does not meet it.
What does an AI red-team audit include?
Adversarial tests on a defined scenario set, not a document review: AIUC-1 runs an agent through jailbreak scenarios (including indirect ones hidden in content it reads), hallucination probes and data-leak tests, at 1,000 to 5,000 scenarios per the standard’s site. Faculty’s scope adds alignment assessments and model-safeguard testing alongside red-teaming.
Where a figure is the company’s own claim or single-sourced, this page says so. AIUC-1 is not regulator-endorsed; no law requires it. On September 20, 2026 Rep. Ro Khanna called for an FDA-style federal AI agency — a television-interview proposal, not law: no agency, no bill text and no approval authority exist, only Congress can create one, as the US AI legislation 2026 page explains.
Sources
- AIUC, “Introducing AIUC-1”, 17 July 2025, and the AIUC-1 standard site — https://aiuc.com/updates/introducing-aiuc-1 · https://aiuc-1.com/
- AIUC Series A release, PR Newswire, 15 September 2026 — https://www.prnewswire.com/news-releases/aiuc-raises-40m-series-a-from-ribbit--first-harmonic-to-build-confidence-infrastructure-for-frontier-ai-302879036.html
- TechCrunch, 15 September 2026 — https://techcrunch.com/2026/09/15/early-anthropic-hire-former-metr-coo-have-found-a-way-to-rein-in-rogue-ai-agents/
- AIUC updates — https://aiuc.com/updates/aiuc-1-certificate-overview
- Anthropic, “Partnering with Accenture on embedded evaluation”, 18 September 2026 — https://www.anthropic.com/news/accenture-embedded-evaluation
- TechCrunch, Tim Fernholz, 18 September 2026 — https://techcrunch.com/2026/09/18/anthropics-first-embedded-evaluator-is-accenture/
- CNBC, Jonathan Vanian, 18 September 2026 — https://www.cnbc.com/2026/09/18/ai-safety-evaluators-anthropic-openai-models-security.html