Key Insights
A senior associate opens an AI-drafted workpaper at 9 p.m. The narrative reads cleanly, the conclusions sound right... and the citations point to procedures that were never performed on this engagement. The manager spends the next two hours rebuilding the analysis against the actual client file and last year's workpapers. The draft wasn't wrong because the model was bad. It was wrong because the model never saw the right material before it answered.
This article covers what context engineering means on an engagement, the four inputs that decide whether agent output is usable, and why dumping the full client file into the model makes the problem worse, not better.
Context engineering is the work of choosing, organizing, and governing the external knowledge, memory, and human input a model sees before it answers. The scope reaches past a single prompt to what the model retrieves and carries forward across a job.
For an audit engagement, context engineering gives the model the engagement material it needs before it reasons:
Without those inputs in front of it, the model is working from training data alone, and the output reflects that.
The catch is that a model can hold only so much at once. What you put in front of it is what it works from; what you leave out, it never sees. Output quality, and the firm's exposure, come down to that selection.
Hand a general-purpose model an audit question with no engagement context and it will still answer. That's the trap. With nothing specific to work from, it leans on its training data and produces audit language that reads well but may have nothing to do with your methodology or this client's facts.
These tools predict the most likely next words from patterns picked up in training, and your client files and methodology aren't in there unless you put them there. So when the engagement material is missing, making things up is a built-in risk, not a malfunction, and the model often won't flag when it's guessing. The versions that sound the most certain are wrong just as often.
For audit work, this often shows up as invented citations. Even advanced models make up citations 39% of the time when they're working without the source material. A workpaper that cites procedures or evidence that don't exist is the kind of error review is meant to catch, and it's more likely when the model never had the source to begin with.
The fix is to put the right evidence in front of the model from the start. When it answers from the actual source documents rather than memory, it makes fewer errors and every conclusion traces back to where it came from.
Better context makes AI more accurate, so it's natural to assume more of it helps. Firms act on that instinct and upload the whole client file, expecting the model to sort out what matters. It doesn't work that way.
Raw volume works against the model. As unrelated documents pile into the context, accuracy slips for every model tested, and even one extra document can drag it down. Long contexts carry a blind spot too: models hold onto the beginning and end of what they're given but miss what sits in the middle. Dump a disorganized file in and the noise crowds out the evidence that matters.
The lesson isn't to use less AI. It's that someone has to choose what the model sees, because poor or unsorted inputs produce untrustworthy results and curating inputs is a required step for output you can trust. That's the difference between dumping a file into a general-purpose model and working in a platform built for the engagement. The platform surfaces the structured, engagement-specific material a step needs: methodology, prior-year work, linked evidence, the framework in scope. The model works from that curated set, not the entire file.
Practitioners who treat context engineering as editorial work get real value from AI; firms that dump files get fluent nonsense. Someone, or some system built on your methodology, has to decide what the agent sees and which inputs to exclude. Firms that treat context as a discipline give reviewers more time judging work and less time redoing it.
A general-purpose model wasn't built for audit work, and no amount of clever prompting changes that. What firms need is AI built for the engagement: purpose-built for audit and advisory, fluent in the methodology, and designed to execute the procedures practitioners actually run.
You can feed engagement context to any model. One built for audit work does more with it, because it already understands how engagements, workpapers, and frameworks fit together. Four inputs decide whether the output holds up on a given job, and missing any one degrades the result in a specific, predictable way.
Without your methodology, an agent defaults to a plausible-sounding approach that may not match how your firm scopes work, sizes samples, or documents conclusions. Fieldguide draws on Agent Knowledge, the firm's methodology and standards, so Field Agents plan, test, and document the way your firm does. Audit programs, control libraries, framework mappings, and review expectations live in one place and flow into every engagement, so the output reflects your approach rather than a generic one.
Last year's files tell the agent what changed and what carried forward. Strip them out and every engagement starts cold. The agent loses the continuity that makes roll-forward efficient and that reviewers expect to see reflected in the workpapers. In Fieldguide, prior-year workpapers, requests, and conclusions roll forward into the current engagement, and Field Agents work from that history when they draft narratives, flag changes, and propose updates to control descriptions and risk assessments.
Linked source documents give a conclusion something real to point to. When the model works from the actual documents under examination, its output can cite them; when it can't, you're back to the plausible-but-wrong answers from earlier. Fieldguide links client-provided evidence directly to the requests, controls, and tests it supports, so Field Agent output carries clickable citations back to the exact page or section a conclusion relies on.
For assurance work, the applicable framework defines what sufficient evidence means. Without it, the agent can't distinguish a complete test from a partial one. The requirements set the bar the work is measured against. Fieldguide ships with the major frameworks built in and lets firms layer their own interpretations and crosswalks on top, so Field Agents test against the specific criteria in scope for the engagement instead of a generic version of the standard.
Incomplete context produces a polished answer that a reviewer then has to take apart. That cost lands on the most senior person on the engagement.
Accurate output is only half of what a reviewer needs. The other half is whether they can confirm it without rebuilding it, and that comes down to whether the output shows its evidence. A conclusion that arrives with its source attached gets reviewed in place; one that doesn't gets reviewed by redoing the work.
What a reviewer is checking for shows up plainly in the regulator's findings. The PCAOB's 2024 inspections flagged 39% of audits where the firm hadn't obtained sufficient evidence to support its opinion, and the recurring problem was workpapers describing procedures that weren't performed or conclusions that didn't tie back to evidence. Confirming that tie is the reviewer's core job, and it's where AI output either saves time or burns it.
When the output carries its own evidence trail, the reviewer's job changes shape. Instead of rebuilding the analysis to confirm it, the reviewer checks the reasoning against the source. Trusting AI output requires understanding which samples and methods produced a given recommendation, and a reviewer needs a record of what happened, how the conclusion was reached, and why it was made.
In Fieldguide, that record is built into every agent run. A Trace shows the inputs the agent used, the output it produced, and the reasoning that connected them. Clickable citations on the document, request, and workspace surfaces point back to the exact source the conclusion relies on.
The result is a different review motion. The reviewer clicks a citation, lands on the highlighted source section, and judges whether the conclusion holds, instead of rebuilding the analysis from the client file.
Fieldguide is the end-to-end AI-native platform purpose-built for audit and advisory, with the methodology depth and audit-grade rigor firms need to run engagement work this way. Context isn't something a team configures after the fact; it's built into how every engagement runs. Field Agents execute against curated engagement inputs, and practitioners review, judge, and own every conclusion. If your reviewers spend more time rebuilding AI output than checking it, the problem sits upstream of the model. AI-assisted output supports professional judgment; it does not replace it. To see how this works on your own methodology, request a demo.