Picture a senior associate starting substantive testing on a revenue balance. That means working through a spreadsheet and a stack of client evidence, one step at a time, with the associate driving every action by hand: pull the population, calculate the sample, request evidence, match each document to each row, extract the figures, compare them against the records, write it up. Each step is a fresh instruction typed into a plain tool built for one job at a time, an Excel formula here, a saved email attachment there, and the associate carries the context between them by hand, re-checking and re-entering the same figures at every switch.
A long-horizon agent changes the starting point. The auditor describes the test once, reviews the plan the agent proposes, and approves it. From there the agent works through the sequence, stopping to check in when it needs a decision or when a step is done. The auditor is still in charge of the judgment. What changes is that they are reviewing work instead of keying it in. This article covers what a long-horizon agent actually is, what it changes for audit teams, how it runs a testing workflow in practice, and why the auditor's expertise matters more in this model, not less.
Most AI tools in audit are assistants. You ask a question, you get an answer, and you decide what to do with it. That is useful for research and document review, but the engagement model underneath it does not move: a person still starts every step, runs it, and decides what comes next.
A long-horizon agent works differently. You give it a goal instead of a single instruction, and it plans the sequence of steps, executes them, and adjusts as it goes, checking in with a person at the points that call for judgment. It holds the thread across the whole workflow rather than forgetting everything between prompts.
Fieldguide splits its AI into two categories. AI Assist covers the on-demand tools a practitioner drives directly, like AI Chat and AI Actions. The Agent Workforce executes multi-step engagement work on its own, with the practitioner reviewing the result. A long-horizon agent belongs to the Agent Workforce: the difference between a tool that helps you test faster and one that runs the test and hands you the result to review.
When one agent carries a test from population through documentation, several things change at once:
The point is not that any one task gets automated. It is that the procedural middle of a test, the part that eats the most hours and requires the least judgment, moves to the agent, and the auditor's time moves to where their expertise is worth the most.
Fieldguide's Field Orchestrator is a working example, built for audit and advisory and used first for end-to-end substantive testing. The auditor opens an audit program, and the testing workspace opens with it. They tell Field Orchestrator what they want to test, and it reads the engagement context and asks for the population file. It proposes a plan, and the auditor adjusts and approves it before any step runs. From there, the agent works through the shape every auditor knows: analyze the population, build the sample, collect and link evidence, extract the values that matter, compare them against what was recorded, and surface any exceptions. It writes the results into the sheet with citations as it goes. The auditor ends up reviewing a finished piece of work instead of assembling one.
The chain of command holds throughout. The auditor directs Field Orchestrator, Field Orchestrator coordinates Field Auditor and its roster of specialized testing agents, and those agents execute the steps. Nobody skips a link, and the auditor still signs off on everything.
As agent output scales, keeping track of it becomes its own job, which is what the kanban-style Field Board solves. It organizes every control and deliverable into a clear lifecycle, so a manager sees engagement status in seconds instead of hunting across separate files. Field Board shows what's happening; Field Orchestrator is how the auditor directs it.
The agent pauses at defined points for the auditor to review the plan, adjust the approach, and approve before execution continues. Exceptions, judgment calls, and elevated-risk items get surfaced for a person rather than resolved silently. Agent-tested controls land in the Agent Review Experience, a single workspace where the reviewer validates the evidence and signs off. The auditor owns every conclusion, every finding, and the final signoff. The agent executes; it does not decide.
Every governance framework in this space converges on the same idea: autonomy needs a place to stop and check. The PCAOB's May 2025 Spotlight found audit committee chairs worried that leaning too hard on automation leaves auditors complacent, exactly the failure mode a checkpoint exists to prevent. The NIST AI Risk Management Framework treats autonomy as a spectrum, with deliberate escalation to a person wherever a system can't catch its own errors. The pattern holds across both: the capability is ready before the guardrails are, and building the guardrails in from the start is what lets an agent earn a place on a real engagement.
Getting this right pays off directly on the clock. Every hour an agent saves on procedural work is an hour a practitioner gets back for review and judgment, the parts of the job that actually need a person. That efficiency only holds up if the review points behind it are real. UHY reported cutting some tasks from three hours to 15 minutes, with agents handling the execution and auditors reviewing the results, time saved without cutting corners on oversight.
Fieldguide gives audit and advisory firms an end-to-end platform where the operating model matches how this work has to run: agents execute the procedural stretches, practitioners review at every checkpoint, and one system carries context from planning through reporting. Because the work is sensitive, the platform is backed by ISO 42001 governance certification and is the first in its category to earn AIUC-1 certification for agentic AI. Practitioners still own professional judgment and final signoff on every output; the platform accelerates the work, it does not replace the judgment behind it. To see a long-horizon agent run a testing workflow on a real engagement, request a demo.