Excel Agents is here - automate analysis and testing inside Excel
See it in action

AI Governance in Audit: Scaling Agents Without Losing Trust

AI & Intelligent AutomationInternal Audit
Blog post featured

Key takeaways 

  • Treat every AI result as a first draft that a qualified auditor reviews and signs.
  • Make every output traceable to the evidence behind it, so a reviewer can open any figure's support in one click.
  • Prove accuracy on last year's completed workpapers before you run an agent on live work.
  • The auditor who signs owns the workpaper, no matter how much AI touched it. PCAOB standards already say so.

One word nearly turned a control-testing agent into a liability. During development, a team changed an instruction from the agent "should" document every attribute to it "must." The agent decided the cleanest way to satisfy "must" was to generate convincing evidence that the test had passed, so it fabricated the workpaper. Luckily, the team caught it in testing, so no harm done.  

That story, from our webinar with Protiviti, is the case for AI governance in audit in miniature. As Angelo Poulikakos, Global Lead, Internal Audit & Financial Advisory, described it, you can get "a fluent, confident, completely wrong result that looks exactly like good work."  

And it isn't hypothetical outside audit. In 2023 a lawyer filed six ChatGPT-generated case citations in federal court, with fake names and quotes, and when opposing counsel questioned them, the tool "confirmed" they were real. A judge sanctioned him (Mata v. Avianca).  
Governance is what catches that before it reaches a client or a regulator. Done right, it's also what earns you the room to move faster. 
 
The full session with Protiviti goes deeper on all of it. Watch it on demand

Governance is the accelerator 

That runs against the instinct that governance is a brake. Andrew Struthers-Kennedy, Global Lead, CAE Solutions at Protiviti flipped it in the session: "Governance isn't to slow things down. It's to help us move fast under control, safely."  

The stakes are what make audit different from marketing or engineering, where leaders can hand people AI and just ask for more output. Trust in the audit function is earned slowly and lost fast, so one wrong conclusion can cost more than any speed gain. Angelo put the balance plainly: move too slow and you risk becoming irrelevant, move too fast with AI and you risk the trust and judgment your team runs on. 

A four-part model for AI governance in audit 

The session turned that principle into four practices, each one a matter of discipline more than tooling. 

Govern the output 

You can't meaningfully inspect how a model reasoned, so put your review on what it produces. In practice that's a defined workflow: the auditor sets the procedure, the agent handles the extraction, matching, analysis, or workpaper prep, and the output comes back with its evidence linked. The reviewer then checks the snips, formulas, exceptions, and conclusions, and the auditor applies judgment and signs off. Nothing goes to sign-off until a person has reviewed it. And the reviewer has to be a qualified auditor who knows what good looks like, because that's who catches the "must" problem. That's also where it gets hard. When we polled the session, team skills and training came out as the biggest barrier to responsible adoption, and 63% named upskilling their current team their top talent concern. Governance leans on skilled reviewers exactly where teams say they're thinnest. 

Make every number traceable to its source 

Every claim, number, and citation should trace back to a piece of evidence a reviewer can open. This is the heart of how DataSnipper works: each output ties to the source behind it, so a reviewer can open the evidence for any figure in one click. Walker & Dunlop's financial reporting team saw what that does to review time: "UpLink helped us turn hours of manual compliance work into minutes. We no longer need to open every file or manually search for figures, everything is searched, snipped, and ready for review. It's significantly reduced both preparation and reviewer time." (Dylan Chaikin, Manager, Financial Reporting.)  

Track error rates from day one 

Trust gets earned with data, so measure how often the agent is right before you rely on it. The method from the session is simple: don't debut an agent on live work. Point it at last year's completed workpapers, where good is already defined, and compare. Move to live engagements only once it clears the bar you set. Protiviti put that bar around 95% or higher, as a practitioner's rule of thumb rather than a formal standard. FEI's framework calls the same technique performance testing: comparing AI output against known "ground truth" results. 

Make ownership explicit 

The auditor who signs the workpaper owns it, no matter how much AI touched it. "This is what the AI generated" will never hold up, and the standards back that up. PCAOB rules put responsibility for the sufficiency and appropriateness of audit evidence on the engagement team, and its 2024 amendments on technology-assisted analysis (effective for fiscal years beginning on or after December 15, 2025) reinforce that as AI enters the workflow. Name who owns each review step, and accountability stays attached to a person. 

How this maps to the frameworks auditors already trust 

Three AI governance frameworks already map this ground, and existing audit rules reinforce them: NIST AI Risk Management FrameworkCOSO's 2026 guidance on generative AI, and FEI's 2026 framework for internal control over financial reporting. Here's how the four practices line up.
Practice 
Where it shows up in the frameworks 
Govern the output 
FEI's human-in-the-loop oversight; NIST's Govern function, which spans the whole AI lifecycle 
Make everything traceable 
COSO's audit-ready control mapping and evidence expectations; PCAOB's evidence requirements 
Track error rates 
NIST's Measure function (evaluate systems and monitor them in production); FEI's performance testing and multi-model validation; COSO's monitoring component 
Make ownership explicit 
PCAOB responsibility for the sufficiency of evidence; FEI's human-in-the-loop oversight 

One point they share is worth keeping in view. All of these treat hallucination, model drift, and bias as risks to manage. No tool has solved them, which is exactly why human review stays at the center. FEI's framework, built by controllers and chief accounting officers from Fortune 100 companies, leads with human-in-the-loop oversight rather than trusting the model to police itself. 

Match the governance to the stage 


Governance scales with how much you're automating. DataSnipper's AI Maturity Model describes the progression in four stages: task-based automation, AI-assisted automation, agentic automation, and connected agents. Early on, a human touches every output, so trust is easier to hold. Further along, the human moves from in the loop to on the loop, overseeing rather than checking each step. 
image (4).png
You earn the right to the next stage by governing the last one. RSM Cayman is a good example of that governed progression, standardizing reviewable audit procedures as it moved deeper into AI-assisted work rather than flipping a switch. Rushing to connected agents before the review muscle is built is how the "should to must" story happens at scale. 

Hear how AI governance gets written, from someone who's shaped it 

Alondra Nelson will speak at DataSnipper Connect NYC. As director of the White House Office of Science and Technology Policy, she led the team behind the 2022 Blueprint for an AI Bill of Rights, and she's a co-author of the 2026 MIT Press book Auditing AI. She was named to TIME's inaugural TIME100 AI list. This piece is about how audit governs AI. Her keynote is about how that governance gets written. 

Want to continue the conversation? Book a demo to see traceable, review-ready AI on your own workpapers.