Key insights
- The search for unrecorded liabilities (SURL) looks for bills the client owed at year-end but never booked.
- Most teams only check the payments made after year-end, missing two spots where those bills usually hide: late-posted invoices and invoices that were never entered at all.
- AI is a good fit for the coverage gap. It can match evidence, rank the population by risk, and flag exceptions, while the practitioner still owns the completeness call.
The search for unrecorded liabilities (SURL) is one of the most common procedures on a financial statement audit, and one of the most quietly flawed. Its job is to catch bills the client owed at year-end but never booked. On most engagements, it still runs the way it has for years: pull a sample of checks written after year-end, tie them back to invoices, and move on. The trouble is that the checks written after year-end are only one place an unrecorded liability can live. Invoices posted late to the payables system and invoices that were never entered at all are the other two, and neither shows up in the check register.
That is a coverage problem, not a judgment problem, which is why AI is a good fit for the fix. This article walks through what the test has to prove, where the manual version breaks down, and how AI can change the shape of the work without changing who signs off on the conclusion.
1. Pin down what the test actually has to prove
The SURL targets one thing: understatement of liabilities at the balance-sheet date. In practice, it is a reciprocal-population procedure. The auditor looks at what happened after year-end (subsequent disbursements, late-posted invoices) to find obligations that should have been recorded before year-end.
Scope follows from that framing. SURL addresses completeness for accounts payable and accrued liabilities, with cutoff riding along: did the obligation exist at the balance-sheet date, and did it land in the right period? Client controls set the temperature.
If the client only records an invoice when they pay it, rather than when they receive it, internal control guidance treats that as a weak control design, because any invoice received but not yet paid at year-end will be missing from the books. The usual audit response is to lower the SURL testing threshold, expand the population, or both.
Skipping the test entirely is the extreme failure mode, and it has happened. In one PCAOB enforcement order, the engagement team's audit program planned a search for unrecorded liabilities but the team never performed it, or any other procedure to test accounts payable completeness.
The more common problem is quieter: teams perform the test on the wrong population. Either failure mode can draw inspection attention. During 2024, 10% of PCAOB-inspected non-traditional focus areas resulted in a deficiency. The PCAOB's staff update tied many of those deficiencies to accruals, debt, equity, and expenses.
2. Find where the manual version breaks
The manual version breaks when convenience shapes the evidence before the testing logic does. Once that happens, later sampling choices can look disciplined while the underlying coverage problem stays unresolved.
The population is bigger than the check run
Ask where an unrecorded liability actually lives at year-end, and the answer is rarely a disbursements journal item waiting to be sampled. A complete SURL goes beyond subsequent disbursements because unpaid purchase or expense invoices may have been entered into the payables system after the balance-sheet date, or may not have been entered anywhere yet. A complete population includes:
- Post-period cash disbursements
- Post-period payables journal entries
- The unentered invoice file (physical or digital invoices the client has received but not yet booked)
When there is a longer-than-usual lag between recording payables and cutting checks (common with financially stressed clients), merging the payables journal and disbursements journal populations, then eliminating duplicates, gives the sample a complete population to draw from. Under deadline pressure, teams often skip that step and test the check run alone.
The test window gets set on autopilot
Teams routinely extend the subsequent-period test through the last day of fieldwork because that is what last year's workpaper did. Extending too far can overstate any projected error and make it unreliable, particularly when exceptions cluster in the early part of the test period (which is where they normally sit). Staffing reality compounds the problem. When scanning replaces sampling in low-risk areas, the work needs someone with enough client knowledge to recognize what is unusual, and that person is rarely the first-year running the workpaper under deadline pressure.
3. Use AI to support matching and population coverage
Those breakdowns are data-coverage and execution problems as much as judgment problems, which is why AI helps here. When dated population data and evidence files are available, AI changes where the team's time goes: less manual tie-out, more exception review tied back to completeness.
AI-assisted analysis narrows the sample-size debate
AI-assisted full-population analysis can make anomaly detection more reliable by working across all transactions rather than a sample. In a SURL, that means risk-ranking the population and focusing review on higher-risk items: entries posted at odd hours, transactions missing approvals, round-dollar amounts where precision is expected, and items that look like year-end obligations paid late.
The matching work can run the moment evidence arrives
Once evidence lands, much of the SURL work looks like pattern recognition: match the disbursement to the invoice, read the service date, compare it to the period under audit, and decide whether the exception needs attention. That layer is where AI creates the most obvious lift on a SURL, and it can be delivered in two very different ways.
Most AI in audit today is human-orchestrated: chat, copilots, and point automations a practitioner triggers step by step. Fieldguide calls this AI Assist and uses it for the parts of the engagement where the practitioner drives each step. Field Agents on the Fieldguide platform sit in a second category, Agent Workforce: purpose-built agents that execute engagement work while practitioners review the results.
For a SURL, Agent Workforce is what changes the day-to-day. Field Auditor gathers and validates evidence as it arrives, executes the defined test procedures, and documents results with citations and exception flags in the workpaper. Agent Triggers start that configured work automatically on document upload, rather than waiting for a staff member to notice the client responded. UHY reported a 20–30% engagement time reduction, with some tasks cut from 3 hours to 15 minutes.
The standards footing is now explicit
Regulators are catching up with what AI-assisted testing can do. For issuer audits, the PCAOB's technology-assisted analysis amendments take effect for fiscal years beginning on or after December 15, 2025. In plain terms, they do two things. First, they confirm that testing the full electronic population (rather than a sample) is an accepted way to gather audit evidence. Second, they raise the bar on how auditors validate the electronic data a client provides, including the information technology (IT) general controls that govern how that data is produced. Nonissuer audits are not directly in scope, but the direction of travel is clear, and the same expectations tend to filter down.
4. Keep the review points that were never mechanical
Field Agent outputs run through practitioner review and approval before finalization. In a SURL, that review is where the output connects back to the actual audit question. Five judgment calls stay with the reviewer no matter how the testing was executed:
- The risk period. Defining when unrecorded liabilities are likely to surface drives whether projected error is reliable at all. Defaulting to the report date can undermine that assessment.
- The threshold. A lower dollar cutoff in response to weak payables controls follows from the auditor's risk assessment, not the prior-year workpaper.
- The unentered invoice question. Whether the client keeps an unentered invoice file, and whether they maintain physical control of it, determines whether extended procedures like payables confirmations are warranted.
- Exception evaluation. Deciding whether a flagged disbursement is a genuine unrecorded liability, a cutoff issue, or a control weakness with fraud implications is the substance of the procedure.
- The conclusion. Accepting or rejecting the difference between expected and recorded balances is a materiality assessment, and it still needs to be documented.
The workpaper is strongest when the basis for each of these five judgment calls is documented alongside the Field Agent output. That is the pattern that connects agent-executed testing to the practitioner's conclusion.
Run the next SURL on Fieldguide
The SURL is a good place to see the new operating model without turning the audit upside down. The repetitive matching moves earlier in the engagement, exceptions surface faster, and the reviewer spends more time on whether the test design supports the conclusion, instead of ticking and tying. Fieldguide gives audit and advisory firms one place for that work across the engagement, with Field Auditor executing test procedures and evidence matching and Field Reviewer surfacing exceptions and higher-risk areas for the human reviewer. Half of the Top 100 firms, including KPMG, already run engagements on the platform. Request a demo and walk through a search for unrecorded liabilities on your own engagement data.