- /
- Blog
AI Governance in Audit: Scaling Agents Without Losing Trust
Key takeaways
- Treat every AI result as a first draft that a qualified auditor reviews and signs.
- Make every output traceable to the evidence behind it, so a reviewer can open any figure's support in one click.
- Prove accuracy on last year's completed workpapers before you run an agent on live work.
- The auditor who signs owns the workpaper, no matter how much AI touched it. PCAOB standards already say so.
One word nearly turned a control-testing agent into a liability. During development, a team changed an instruction from the agent "should" document every attribute to it "must." The agent decided the cleanest way to satisfy "must" was to generate convincing evidence that the test had passed, so it fabricated the workpaper. Luckily, the team caught it in testing, so no harm done.
That story, from our webinar with Protiviti, is the case for AI governance in audit in miniature. As Angelo Poulikakos, Global Lead, Internal Audit & Financial Advisory, described it, you can get "a fluent, confident, completely wrong result that looks exactly like good work."
The full session with Protiviti goes deeper on all of it. Watch it on demand.
Governance is the accelerator
That runs against the instinct that governance is a brake. Andrew Struthers-Kennedy, Global Lead, CAE Solutions at Protiviti flipped it in the session: "Governance isn't to slow things down. It's to help us move fast under control, safely."
The stakes are what make audit different from marketing or engineering, where leaders can hand people AI and just ask for more output. Trust in the audit function is earned slowly and lost fast, so one wrong conclusion can cost more than any speed gain. Angelo put the balance plainly: move too slow and you risk becoming irrelevant, move too fast with AI and you risk the trust and judgment your team runs on.
A four-part model for AI governance in audit
The session turned that principle into four practices, each one a matter of discipline more than tooling.
Govern the output
You can't meaningfully inspect how a model reasoned, so put your review on what it produces. In practice that's a defined workflow: the auditor sets the procedure, the agent handles the extraction, matching, analysis, or workpaper prep, and the output comes back with its evidence linked. The reviewer then checks the snips, formulas, exceptions, and conclusions, and the auditor applies judgment and signs off. Nothing goes to sign-off until a person has reviewed it. And the reviewer has to be a qualified auditor who knows what good looks like, because that's who catches the "must" problem. That's also where it gets hard. When we polled the session, team skills and training came out as the biggest barrier to responsible adoption, and 63% named upskilling their current team their top talent concern. Governance leans on skilled reviewers exactly where teams say they're thinnest.
Make every number traceable to its source
Track error rates from day one
Trust gets earned with data, so measure how often the agent is right before you rely on it. The method from the session is simple: don't debut an agent on live work. Point it at last year's completed workpapers, where good is already defined, and compare. Move to live engagements only once it clears the bar you set. Protiviti put that bar around 95% or higher, as a practitioner's rule of thumb rather than a formal standard. FEI's framework calls the same technique performance testing: comparing AI output against known "ground truth" results.
Make ownership explicit
How this maps to the frameworks auditors already trust
Practice | Where it shows up in the frameworks |
Govern the output | FEI's human-in-the-loop oversight; NIST's Govern function, which spans the whole AI lifecycle |
Make everything traceable | COSO's audit-ready control mapping and evidence expectations; PCAOB's evidence requirements |
Track error rates | NIST's Measure function (evaluate systems and monitor them in production); FEI's performance testing and multi-model validation; COSO's monitoring component |
Make ownership explicit | PCAOB responsibility for the sufficiency of evidence; FEI's human-in-the-loop oversight |
One point they share is worth keeping in view. All of these treat hallucination, model drift, and bias as risks to manage. No tool has solved them, which is exactly why human review stays at the center. FEI's framework, built by controllers and chief accounting officers from Fortune 100 companies, leads with human-in-the-loop oversight rather than trusting the model to police itself.
Match the governance to the stage
.png)
Hear how AI governance gets written, from someone who's shaped it
Alondra Nelson will speak at DataSnipper Connect NYC. As director of the White House Office of Science and Technology Policy, she led the team behind the 2022 Blueprint for an AI Bill of Rights, and she's a co-author of the 2026 MIT Press book Auditing AI. She was named to TIME's inaugural TIME100 AI list. This piece is about how audit governs AI. Her keynote is about how that governance gets written.

