Resource Articles

AI agents in audit: making the exception record defensible

Written by Amanda Waldmann | Jul 21, 2026 9:58:26 AM

Key Insights

  • When AI agents test overnight, the manager's morning shifts from performing work to adjudicating exceptions.
  • But there's a right way and a wrong way to clear an overnight queue of agent-tested results. A wall of green checkmarks with one click-through approval is what a PCAOB inspector flags as a Part I.A deficiency.
  • Agent reasoning doesn't leave a trail by default the way a spreadsheet does. Unless the system is built to capture it, the file loses the basis for the conclusion by the time anyone reviews it.

A manager walks in Tuesday morning to a queue of 400 test procedures an agent ran overnight, every row green and the sign-off one click away. Clearing that queue is easy. Building a record that shows an inspector why each result was trusted is the harder part, and getting it right starts with choosing a platform built for that job. This article covers what substantive review looks like when agents execute the work, why agent reasoning takes more effort to document than spreadsheet work, and what a record needs to show to hold up under inspection.

Clearing the queue isn't the same as reviewing it

AI agents now do the testing that used to fill an associate's week: pulling the population, matching evidence to samples, running the procedure, drafting the workpaper narrative. The work still gets done. It just shows up differently: finished, in bulk, by morning.

Clearing that queue is not the same as reviewing the work. Human oversight is something firms have to measure and manage, because a person in the loop does not automatically mean the risk is covered. Real oversight means examining what went into the agent and judging whether what came out is reliable enough to rely on, before signing off, not after.

Skip that step and automation bias takes over. Commission bias means accepting an agent's result despite evidence that contradicts it. Omission bias means failing to act because the agent didn't flag anything. Both failures land on the person who cleared the queue: when the agent does more of the work, the review has to do more too, or a control that always passes stops being a control.

How the manager's role shifts from production to adjudication

When an agent has already gathered the evidence, tied out the population, run the procedure, and drafted the workpaper narrative, re-performing that work doesn't add anything. The manager's job shifts from doing the work to judging it: deciding which outputs to trust, and pressure-testing the conclusions that matter, even the ones that look clean.

Judging the work is a different muscle than reviewing a junior's. A senior used to catch a bad tie-out by re-footing the schedule. Now the schedule already ties, so the real questions move up a level: did the agent pull the right population, does the sample it selected actually address the risk, and does the exception it flagged as immaterial deserve a second look. Young professionals need that same instinct now, learning to evaluate and challenge AI-generated output the way they once learned to challenge a colleague's draft. The standards have not moved: accuracy, judgment, due care, and professional skepticism still sit with the practitioner, whether AI drafted the work or automated it outright.

In practice, that looks like this: instead of re-performing 60 control tests, a manager spends the morning on the four that came back with exceptions and the two revenue cutoff calls that need judgment. That is the senior-level work firms already value. The production underneath it is agent-executed now, so the hours land on the part of the engagement that actually needs a person.

What makes agent-executed review defensible

Review only works if there's something to check it against. For agent-executed work, that means the file has to show why each result came out the way it did, including any pushback or escalation along the way. Without that, a reviewer is left trusting memory and gut feel instead of evidence.

Why agent reasoning is harder to trace than a spreadsheet

A spreadsheet shows its formulas. Click a cell and the logic is right there. An agent's path to a conclusion only appears when the system was built to capture it.

That is because an agent's decision path branches, shifts internal state, and runs through several transformation steps before landing on an answer, and conventional system logs were not built to reconstruct any of that after the fact for compliance or root-cause work. Knowing the agent finished a step is not the same as knowing why it decided the way it did.

The reviewable record has to be built into the process, not reconstructed after. The file needs to show the degree of oversight applied and tie any overrides to the adjudication and go/no-go decisions behind them. That gives a reviewer a path through the judgment instead of a list of system events.

How the record connects inputs to outputs

A usable work surface shows three things: what the agent took in, what it produced, and the reasoning that connects the two, with the source documents one click away. That's what visibility and traceability actually look like in practice: every state change and approval tied to the work itself instead of scattered across systems.

That kind of record turns "I looked at it" into something a second person can actually verify. The evidence lives with the conclusion instead of getting buried in email threads and shared drives, which is what makes exception adjudication a realistic way to spend a manager's morning.

The record has to hold up under PCAOB inspection

A defensible engagement needs documentation showing that you evaluated the work. A system-run trail alone can lead to a deficiency finding, and the inspection record points to that gap.

PCAOB documentation standards make the point plainly: audit documentation is the written record behind the auditor's conclusions, clear enough that an experienced auditor with no prior connection to the engagement can understand what was done and why. The standard also puts the burden of proof on the auditor: if it later looks like a procedure wasn't performed, the auditor has to demonstrate with persuasive evidence that it was, that the evidence was obtained, and that the conclusions were sound. A click-through approval has nothing to demonstrate with.

A sign-off alone doesn't meet that bar. It shows who did the work and who reviewed it, but not the nature of the work, the results, or the judgment calls behind it, and oral explanation doesn't count as documentation. For agent-executed work, that means the file has to name the specific work that was evaluated, the evidence behind the conclusion, and the reasoning applied to any exceptions.

How the work changes when the record holds up

The clearest way to see the shift is to compare before and after. Before, a senior spent the morning re-footing schedules, tying evidence to samples, and cleaning up an associate's workpaper before the manager saw it. With Fieldguide's Field Agents running that work overnight and logging a Trace of every input, output, and exception along the way, the workpaper arrives already assembled, tied, and referenced. The senior opens it looking for the two sample selections that do not sit right and the exception the agent flagged as immaterial that probably is not, because the record already shows exactly where to look.

The manager's day changes the same way. Less time chasing status across three tools and re-performing junior-level work. More time on the four exceptions from overnight testing, the revenue cutoff call the client pushed back on, and the two areas where the risk assessment needs a second look before fieldwork closes. Firms running Fieldguide's Agent Workforce already report the same shift: review starts at the conclusions, not at the file build, because the Trace behind each result is already there to check.

That is a better use of experienced time, and it concentrates responsibility. The control layer becomes the place where quality is won or lost, and the manager or senior is the person carrying it. The tedium that drives burnout comes down. What remains asks for real skepticism, and with a defensible record already in place, that skepticism has something solid to work from instead of a wall of green checkmarks.

How Fieldguide makes the exception record defensible

Fieldguide is built for exactly this operating model: practitioners and Field Agents on every engagement, agent-executed and human-reviewed. Every agent run produces a Trace connecting inputs, outputs, and reasoning, so the reviewable record exists as the work happens instead of getting rebuilt at review time. Fieldguide holds ISO 42001 certification for AI governance and was the first audit and advisory platform to earn AIUC-1 certification for agentic AI security, safety, and reliability. Field Agents execute the testing and documentation; practitioners still review, judge, and sign off on the results. Request a demo to see how the review record holds together on a live engagement.