Key insights
- "A human stays in the loop" costs nothing to say. The claim matters only when the named reviewer can trace the output and reject it.
- Reviewing finished-looking AI output is harder than preparing the work yourself. Automation bias lives in passive monitoring.
- PCAOB standards do not yet set AI-specific review requirements, and inspectors are already asking about technology use.
Ask any firm how they're handling AI risk and you'll get the same four words back: a human reviews it. It's the sentence that ends every partner meeting, every RFP response, every vendor demo. And it settles nothing. The review behind those four words is where audit quality, inspection readiness, and client trust actually get decided, and the firms that have built it are already pulling ahead on all three.
What "human-in-the-loop AI" actually means
The phrase sounds specific. It isn't. The NIST AI RMF puts human control on a spectrum, from systems running inside guardrails to systems where an expert makes the call. "A human is in the loop" can describe any point on that line. Same phrase, very different work.
That's why most firm AI policies read like reassurance instead of governance. Real oversight names the roles. Who monitors the system. Who pushes back when the output looks wrong. Who reviews before it goes out. Who escalates when a result doesn't hold up. Skip that step and the policy protects nothing.
Client-facing AI output has to pass through the same supervision your firm already has. Tax filings, opinions on financial statements, and anything else that carries the firm's name shouldn't leave the building with an unsupervised AI conclusion inside it. That's the whole point of the supervision policies you already run: they now have to cover a new kind of preparer.
Here's the test. A human-in-the-loop claim only holds if the named reviewer has the evidence, the time, and the authority to reject the output. Take away any of the three and the reviewer is carrying accountability with nothing to back it up.
Human-in-the-loop vs. on-the-loop vs. in-command
Three phrases, three different jobs.
- In the loop: a person signs off inside every decision cycle. Fine for high-stakes calls. A disaster if you apply it to everything, because senior hours drain into work that didn't need them.
- On the loop: you set the guardrails at design time and watch the runs. Good for volume work with clear rules.
- In command: the firm-level call about whether AI touches a given piece of work at all. This one lives with leadership, not the engagement team.
For engagement work, match the model to the risk. Low-risk admin work runs on-the-loop with spot checks. Work that cites authority or supports a revenue cutoff conclusion needs a reviewer in the loop, because the evidence and sign-off matter. The lower the practical oversight, the harder the firm has to work on system testing and governance.
A blanket "a partner signs off on everything" posture sounds safe and doesn't work. It burns senior hours on admin tasks. It also stretches the same partner too thin to catch the calls that actually needed them. The firms getting this right run all three models at once, on different work.
Why review is the hard part
Across financial services, talent is already shifting from doing the work to reviewing the work. Analysts turn into agent managers. Preparers turn into reviewers. Banking is a few years ahead, and the pattern is clear: the industry keeps treating review as the easier half of that trade. It isn't. And the moment deadline pressure hits, the cracks show.
Automation bias is the default
Think about the last time you reviewed a workpaper someone else prepared, versus the last time you built one from scratch. Different job, different brain. Building forces you to touch every piece of evidence. Reviewing means scanning something that already looks finished.
Now hand that reviewer an AI-generated workpaper. Formatted, cited, tied out, complete. Looking done is exactly what generative output is best at.
The profession already has a name for what happens next. Automation bias is the instinct to trust output that came from a machine, especially when it looks polished. The firm that signed the opinion carries the exposure. The vendor that produced the draft does not.
Over-trust and under-trust both cost you
Being a great auditor doesn't automatically make you a great reviewer of AI. Different skill. Good AI review design means knowing where the system tends to fail, where bias creeps in, and where its confidence outruns its evidence. Without that, reviewers can't tell when to push back.
The two failure modes both hit the P&L:
- Over-trust: the reviewer signs off on output that shouldn't have passed. Quality slips.
- Under-trust: the senior quietly re-performs every sample the agent already tested. You paid for the license and kept every manual hour.
The skill in the middle, calibrated review scaled to the risk of the procedure, is what firms actually need to build.
What good review looks like when agents do the work
Here's the thing nobody gets a pass on: the standards don't change because an agent did the work. The engagement partner still owns the result, full stop. PCAOB supervision requirements keep that responsibility inside the firm. The file still has to show that significant judgments were reviewed and that the conclusion ties back to sufficient appropriate audit evidence.
None of that shifts when the preparer isn't a person. Under PCAOB documentation requirements, your file still names who performed and reviewed the work, with dates. For agent-executed work, it also has to explain what the agent did and support the judgment with evidence.
For AI summaries of policy documents or draft findings, the review has to leave a paper trail. The file needs documented review controls: the prompt, the source document, and the reviewer's recorded judgment.
Real review = checking the draft against the source. That's it. That's the whole line between review and rubber-stamping. A memo readability check isn't review. That's proofreading.
What that looks like on a Tuesday in March
A manager opens their laptop and sees a review queue instead of a stack of files. The flagged exceptions come first. For each one, they open the cited source document and decide whether the evidence supports the agent's conclusion:
- A revenue cutoff exception? Pull up the shipping documents and check the dates.
- A sample that tied cleanly? Spot-check it.
- Shipping document dated January, tied to a December invoice? That's a real exception. Note what you checked, send it back.
- Timing gap the contract's shipping terms already explain? Write down the reasoning, close it, move on.
Read the file cover to cover in page order instead, and the budget's gone before you hit the items that actually needed judgment.
That kind of review only works when the source trail sits next to the work. On Fieldguide, every agent run produces a Trace showing the inputs, the outputs, and the reasoning between them. Field Reviewer pulls the exceptions and judgment calls to the front, with elevated-risk items surfaced first. The manager's attention lands on the things that need it, instead of hunting through a clean-looking file for something to question.
Governance turns policy into evidence
Here's the split showing up in AICPA & CIMA research: 88% of senior finance and accounting leaders think AI will be the biggest technology shift of the next 12 to 24 months. Only 8% think their own organization is ready for it. That's not a gap. That's a chasm.
Firms are trying to close it by writing policies. Good instinct, wrong finish line. A policy that lives in a PDF on the intranet does nothing. What matters is whether the file proves the policy actually operated on a specific engagement.
Regulators aren't waiting around either. PCAOB standards don't yet include binding AI-specific review requirements, and the board's Data and Technology project is still working out whether changes are needed. Meanwhile, inspectors are already asking how auditors assess a public company's use of AI, per the 2024 inspection activities spotlight. Translation: no perfect rulebook, plenty of scrutiny.
And that scrutiny is landing in a market that isn't ready. Only one in five companies have mature governance for autonomous AI agents, according to Deloitte research. In a profession that sells assurance for a living, that gap isn't a problem. It's an opening.
So what does mature governance actually look like on a file? A stranger should be able to open the engagement record and follow it end to end:
- Who reviewed the work, and when they signed off
- The agent's prompts or instructions
- The source-check trail behind each output
- Every exception, including the ones closed without adjustment, with the reasoning attached
Nothing exotic. Your documentation obligations already cover this. You're just extending the same file to a nonhuman preparer.
The upside is where this gets interesting. When your firm can walk a client stakeholder or an inspector through exactly how an agent's output got validated on a specific engagement, you're showing something four out of five companies can't. Governance stops being a compliance line item and starts being a reason a client picks you.
Build review into the work
The engagement lifecycle runs in one system, so the work stays connected to the evidence. Practitioners direct the work through Field Orchestrator, and Field Orchestrator coordinates the Field Agents across the engagement. Practitioners review each output and own the professional judgment.
Agent Review Experience is a dedicated workspace for inspecting and appending to agent output before it moves up the chain. Fieldguide holds AIUC-1 and maintains security certifications and AI governance attestations. Request a demo to walk through a live engagement.