Key insights
- Speeding up individual tasks doesn't get the engagement done any sooner. The hours that stretch it out live between the steps, not inside them, and a task tool can't see them.
- Engagement work builds on itself. Every step picks up where the last one left off, so a tool that forgets overnight hands the cleanup back to your team every morning.
- Persistent memory, governed autonomy, and drift monitoring aren't fine print. They're the choices that decide whether AI shows up in your margins or just makes a few tasks faster.
Your firm rolled out AI this year. Extraction is faster, drafting is faster, tie-outs are faster. And the engagement still closed on the same date it always closes. If that sounds familiar, you're on one side of a line the profession is now splitting along. Fieldguide's 2026 survey of 400 audit and advisory leaders found the market has divided in two: 51% are active deployers, running AI across most or all of their engagements, and the other 49% are casual users, still running it ad hoc, one task at a time. Both groups bought AI. The deployers are pulling ahead, and not because they bought better tools. It's because they run AI across the whole engagement, not just the tasks inside it. That's what this article is about: what changes when AI runs on the engagement's clock instead of the task's, and what a firm has to build to turn that into capacity.
What "long-running" actually means for engagement work
A long-running agent holds one goal across a stretch of work and runs the steps to get there. It plans, executes, catches its own mistakes, and adjusts as the engagement shifts underneath it, over dozens of steps, not one prompt. That's the bar Fieldguide built its Agent Workforce to clear, and it's the bar any vendor using the word "agent" should have to clear too.
Maintained state is what separates agents from copilots. A copilot that drafts a memo when you ask has no state to manage: the request comes in, the output goes out, and the next request starts from zero. An agent that picks up evidence in the morning, tests it in the afternoon, and flags an exception two days later is carrying context the whole time. When materiality gets revised on day nine, the agent knows which samples were selected against the old number and which conclusions have to be reopened.
Real agents need five things: autonomy, context, tool use, memory, and a goal they can adjust on the fly. For engagement work, that means moving the file forward without losing your firm's method, staying tied to the risk assessment instead of the average of the internet, and working inside the workflow instead of narrating from outside it. Miss any one and you've got an assistant with a longer marketing description.
Here's the quick test: turn the tool off overnight and ask what it remembers in the morning. Nothing? Then it's task-level automation, whatever the datasheet says. Agent washing is all too common right now: slapping the word "agent" on a chatbot, an RPA script, or an assistant that has none of the five. Don't give the word agent any credit until the memory question gets answered.
Where task-level automation runs out of room
Task automation delivers real gains at the task level. Extraction, tie-outs, drafting a request list, all of it runs faster now. But faster tasks don't add up to a faster engagement, because the tasks were never where the time went. Speed up one step and the wait just moves to the next one: quicker sample selection makes evidence review the new bottleneck, and clearing that pushes the delay down to documentation.
If you've run multiple engagements at once, you know where the hours actually go. Chasing status. Re-explaining the file to whoever picks it up next. Reconciling testing findings that no longer line up with the planning memo. The hours between tasks add up faster than the hours inside them.
Those between-task hours have a name: coordination cost. It's the time a file sits in a queue waiting for the next person, and the meeting where three people re-align on what planning meant. A task tool works inside a step. Coordination is the space between steps, so speeding up tasks does nothing for it, and that space is where most of the engagement lives.
So the fix isn't more task tools. The common move is to bolt AI onto a broken process and run it faster, which makes each step quicker and changes nothing about how the firm works. The move that pays off is redesigning the work itself. Build the firm, not the platform.
Why engagement work is a long-running problem by nature
Engagement work runs for weeks because the standards make it run for weeks. PCAOB's planning standard treats planning as a continual process that keeps going until the audit is finished, not a step you wrap up in week one. The context an agent needs on day 20 was set on day two and revised along the way. A tool that remembers nothing from day one can't carry any of it.
Evidence works the same way. A control that tests clean lets the team pull back substantive procedures; a failure later expands them. If materiality drops mid-engagement, the risk assessment and planned procedures all have to be reworked against the new number. Every phase builds on the state the phase before it left, which is what stateful means, and a stateless tool just hands that reconciliation back to the practitioner.
And it all funnels into one gate. Under the amended AS 1215, effective December 2026, the documentation and supervisory reviews have to be complete before the report release date, and the engagement quality reviewer signs off only after that. A review that turns up problems late can blow the deadline even when the work is basically done. That's the shape of an engagement: weeks of interdependent workstreams converging on a single date. It's the opposite of a task, and a tool that sees one task at a time doesn't even know the gate is there.
What has to hold for an agent to run across days
Three architectural choices decide whether an agent can survive this runtime. Miss any of them and the tool falls back to task-level, whatever the label says.
Persistent state
The agent has to carry the risk assessment, the sample selections, the working conclusions, and the open exception list from one session to the next, without a practitioner re-priming it every morning. It also has to know who did what and when. This is an architecture decision made before the first model call. Most relabeled assistants never made it.
Governed autonomy
Governed autonomy means agent behavior is bounded by policy and permissions built into the workflow, with escalation paths inside the system. The longer an agent runs, the more decisions it makes while nobody's watching. A permission that only lives in a written policy is one the agent will find a way around at 2 a.m. on day seven. Firms that can show a client exactly how the agent was constrained, what it escalated, and what a human approved close deals other firms don't.
Your diligence checklist should test ownership first: who's responsible when too many agents run at once? It should test whether autonomy can drift past its original permissions, and reviewers should see what the system did before it ever touches a live engagement.
Drift monitoring
Adaptive monitoring tracks whether the model or its inputs are drifting from expectation. A short-running tool has limited drift exposure. One that runs across an engagement has more: a client's document format changes mid-fieldwork, an assumption from day two stops holding on day 12, and the agent's outputs slowly stop matching the firm's method. Nobody notices until review, when the pattern is already through the workpapers. Drift monitoring isn't compliance overhead; it's what keeps the capacity gain real.
Security scales with runtime for the same reason. Many AI agents remain vulnerable to hijacking when an attacker slips malicious instructions into a document the agent reads. An agent ingesting client files for weeks gives an attacker far more surface area than a one-shot extraction tool. That's why agent-level certification and a full audit trail belong on your checklist.
What this changes about firm capacity
Look back at those two groups from Fieldguide's survey and the gap shows up right where it counts. Among active deployers, 74.5% report profitability up 10% or more, against 52% of casual users, and 70.1% are taking on more engagements with the same headcount, against 46.9%. The figures are self-reported and describe correlation, not causation. But the thing separating the groups isn't the tooling. It's how far the operating model was rebuilt around it, and the gap compounds every cycle the casual users wait.
When agents handle the execution, the bottleneck in your practice moves. It's no longer how fast your team can grind through the work; it's how much senior attention you have for review, judgment, and the client conversation. Your best people spend their time on exceptions instead of production. Capacity comes from that shift, not from the tools themselves, which is why you have to redesign the practice and not just install AI on top of it.
Concretely, that means giving your senior reviewers a live view of what needs them: the exceptions, the aging requests, the documentation risk building toward the release gate. That's where their attention actually changes the outcome. Rebuild your review paths, escalation rules, and client updates around the agents, and AI turns into real capacity. Leave the old queues in place, and all you've bought is faster tasks and the same margin as last year.
And that capacity is for growth, not cost-cutting. You can't hire your way out of an unstable talent pipeline, so the hours come back as more engagements and deeper advisory work, with partners spending time on business development instead of clearing workpaper comments. Firms that treat agents as a headcount play are missing what the capacity is actually for.
Where checkpoints fit in a longer-running workflow
Long-running workflows still need control points. A checkpoint is a defined moment where the agent pauses, a practitioner reviews the state of the work, and the workflow either continues or gets redirected. Where you put them is a real tradeoff. Too many and the coordination cost the agent layer was supposed to remove comes back. Too few and drift has time to propagate across a dozen samples before anyone looks.
High-judgment moments (a change in risk assessment, an exception disposition, a working conclusion) belong at a checkpoint, and the agent never takes those on its own. Repetitive execution can run longer between checkpoints, because the pattern is stable. In practice, that means checking the work before testing begins, as evidence lands, and at a final documentation gate before release. Get that placement right and the checkpoint layer stops being overhead. It's what makes the whole model safe to run.
That's the model Fieldguide is built on: practitioner led and agent executed. Field Agents execute the work, and every run produces a Trace showing its inputs, outputs, and reasoning, so the practitioner reviewing at each checkpoint can see exactly what the agent did. For associates, it's the more interesting version of the job: they test whether the work holds up while the agent handles the data movement between files.
Run the whole engagement, not just the tasks
Fieldguide is the end-to-end AI-native platform purpose-built for audit and advisory firms. It runs your engagements on a single platform from scoping through reporting, with agents executing across days and practitioners reviewing at the checkpoints that matter. In 2026, Fieldguide will have a full set of Field Agents across the entire engagement workflow, as well as supporting analytics and orchestration, live in production.
Fieldguide was the first AI platform for audit and advisory to achieve AIUC-1 certification. Half of the top 100 firms, including 8 of the top 10, already run engagements on it. See how firms like yours made the shift in the case studies, or request a demo to watch Field Agents work a real engagement file.