Key insights
- Vouch every recorded invoice and tie out every tick mark, and you can still miss the misstatement: the omitted expense sits outside the population you sampled.
- Sample size only buys precision. Direction of testing decides which assertion you're answering, and occurrence and completeness never share one sample.
- Raise completeness risk and the test moves outside the ledger to subsequent disbursements, where most of the work is document handling, not judgment.
Picture an AP workpaper open on a manager's screen at review: 40 vouched invoices, every one tied to a purchase order, receiver, and vendor bill, every tick mark green. The file looks finished. What it doesn't address is the invoice sitting in someone's inbox that never made it into the ledger, or the December bill coded to January to protect the margin. That's where expense misstatement usually lives, and a workpaper built entirely from recorded items can't see it.
Get the direction wrong and sample size can't save you: a bigger sample just measures the wrong population more precisely. Direction is the call that decides whether the workpaper tests the assertion that actually carries the risk on an expense engagement.
Start with the risk: Expenses get understated
Here's the counterintuitive part. The instinct on an expense account is to check that recorded charges are real: catch the padded invoice, the cost that shouldn't be there. That's testing for overstatement, and it's where most occurrence work goes. But management pressure almost always runs the other way. A CFO staring down a covenant breach, or with a bonus riding on the numbers, wants income to look better, and the way to get there is to keep expenses off the books. Every invoice left in a drawer or slid into next year is an unrecorded expense, and an unrecorded liability with it.
The enforcement record backs that up. In December 2024, the SEC charged Becton Dickinson for overstating income by failing to record the cost of fixing known software flaws in its Alaris infusion pump. Leaving roughly $50 million in probable costs off the books overstated fourth-quarter operating income by 82%, and the company paid a $175 million penalty to settle. The costs that should have been on the ledger simply weren't.
That's why direction matters. A workpaper built only from recorded invoices, however tidy, confirms the charges you can see are real. It tells you nothing about the ones that never made it onto the books. Push harder on what might be missing when any of these show up on the engagement:
- The client is bumping up against a profitability covenant.
- Management has a performance bonus on the line.
- Expenses came in below your expectation and nobody has a clean explanation.
- A sale of the business is pending.
When those flags are on the table, completeness is where the engagement risk lives, and the testing hours should live there too.
The five assertions expense testing has to cover
Five assertions do the work on expenses, and they don't all pull in the same direction:
- Occurrence: did the recorded charge actually happen, and does it belong to the entity?
- Completeness: what should be on the books but isn't?
- Accuracy: does the math hold up?
- Cutoff: did the expense land in the right period?
- Classification: did it hit the right account?
Occurrence and completeness send you looking in opposite places. Occurrence starts with a recorded charge and works back to prove it happened; completeness starts from what might be missing and has to look outside the ledger. One sample can't cover both. SAS No. 145 makes that choice more deliberate: pin down which assertion actually carries the risk, and put the testing hours there instead of spreading them across all five. Under the standard, inherent risk and control risk are assessed separately, and when you don't test controls, the risk of material misstatement is just the inherent risk. The assertions worth testing get identified from that inherent risk, before controls enter the picture.
The practical effect: firms are testing fewer assertions and going deeper on the ones that matter. First-year adopters cut the number of assertions they treated as relevant and concentrated the work where the real risk was.
For most AP and operating expense populations, that means completeness lands at high risk while occurrence and cutoff sit at moderate. Spreading equal effort across all five is how a budget gets spent proving things nobody was worried about in the first place.
How to test occurrence: Vouch to the source
Occurrence testing is the direction most auditors default to: pick entries out of the expense ledger or the cash disbursements journal and chase each one backward to the paper that proves it's real. Four questions do most of the work:
- Does a vendor invoice support the amount?
- Does a receiving report confirm the goods or services arrived?
- Does a purchase order show the transaction was authorized?
- Does the disbursement tie to the bank statement?
That's the three-way match: ordered, received, billed. When there's no receiving report (think consulting fees, SaaS subscriptions, legal work), the contract stands in for it, and you're confirming the invoiced amount lines up with the rate the parties agreed to.
One trap worth flagging: an internal voucher isn't the same as a vendor invoice. Under AU-C 500, evidence from independent outside sources is more reliable than anything the client generated for itself. Build the file on internal documents alone and the occurrence conclusion is thinner than the sample size makes it look. Note where each selection came from and how you picked it, so a reviewer can retrace the logic without guessing.
How to test completeness: Search for what's missing
Completeness testing hunts for expenses that should be on the books but aren't. And by definition, omitted items don't live in the ledger, so the sample has to come from somewhere else: subsequent cash disbursements. That's the whole idea behind the search for unrecorded liabilities, the closest thing auditing has to a standard completeness test.
Two design calls decide whether the search catches anything at all.
The population
Start with payments made after year-end, then add unpaid invoices sitting in AP or on someone's desk. Trace anything received before year-end into payables. Confirm anything received after year-end stayed out. Sample only the paid disbursements and you'll walk right past the unpaid invoice, which is exactly where the unrecorded liability tends to hide.
The window and the threshold
Findings normally concentrate early in the subsequent period, so pick a window that reflects where unrecorded items actually surface instead of dragging the test to the last day of fieldwork. Overextending doesn't strengthen the conclusion; it inflates projected error and eats hours.
As assessed completeness risk goes up, drop the dollar threshold on the payments you vouch. And if a handful of large vendors account for most of the year's spend, don't sample at all: reconcile recorded payables directly to vendor statements or confirmations. That's often the faster and stronger procedure.
How to test cutoff: Check the period boundary
Timing games happen at period end because that's the last window where an adjusting entry can still move the number. Inspect invoices and contracts booked around year-end and match them to when the goods or services actually showed up. Check accruals and prepaids for expenses paid in one period but incurred in another. And you can piggyback on work you've already done: the subsequent-disbursements sample you built for the unrecorded liabilities search covers most of the cutoff ground too. One selection, two procedures.
How to test accuracy: Recompute and reconcile
Accuracy failures are arithmetic and posting slips, and the fixes are just as mechanical:
- Recompute extensions and footings on sampled invoices.
- Recompute discount treatment on sampled invoices.
- Agree posted amounts back to the invoice.
- Reconcile the payables subledger to the general ledger.
Vendor statement reconciliations earn their spot on this list too. They catch duplicate payments and misposted amounts the client's own records won't flag, because some expense errors never touch an expense account at all.
How to test classification: Watch the capitalize line
Classification gets expensive at one specific line: capitalize versus expense. Booking something as an asset that should have been expensed lifts income just as effectively as leaving the invoice off the books. Advertising costs parked on the balance sheet are a classic version, and it still turns up in restatements. Test both sides of the line:
- Vouch fixed asset additions to source support.
- Confirm repairs and maintenance didn't get capitalized.
- Investigate debits to property accounts that didn't come from actually acquiring anything physical.
- Scan expense accounts for postings that belong somewhere else.
Sampling: Let the assessed risk set the scope
Sample size isn't a number you look up. Risk sets it. Higher assessed risk of misstatement and higher expected misstatement both push the sample up; lower tolerable misstatement does the same. And on a large population, the size of the population itself barely moves the number, which surprises people every year.
Think in two tiers. First, examine 100% of the items where accepting sampling risk isn't a defensible choice: anything that could, on its own, blow through tolerable misstatement. Then look at what's left. Sometimes that's a sample. Sometimes a tight analytic on a predictable account (rent, some utilities) does the job. For significant risks, though, tests of details carry the weight; analytics alone are unlikely to be sufficient.
Nonstatistical sampling is still on the table, but the size still has to reflect the risk of incorrect acceptance: the risk you sign off on a balance that's materially wrong. PCAOB inspections keep flagging insufficient testing of a significant account or identified risk as a top recurring deficiency. Scope, not methodology, is where inspectors bite hardest.
Where the hours actually go
Pull up any expense workpaper and the shape of the work is always the same:
- Locating invoices.
- Extracting amounts and dates.
- Matching documents to selections.
- Tying totals.
- Formatting the tab.
None of that is judgment; it's document handling. And it's exactly the work now being automated. KPMG has put search and data extraction in scope for an 80% time reduction in its latest audit-quality report, and those two tasks cover most of the list above.
Fieldguide's Field Agents are built for exactly this layer: they pull the evidence, run the test across the sample, and draft the workpaper. Your senior reviews the result instead of building it from scratch.
Run expense testing where the evidence lives
Expense testing runs on documents, and it goes faster when the evidence sits next to the workpaper instead of in a separate system.
That's the point of a single platform. Fieldguide's financial audit platform covers the engagement end to end, from risk assessment and analytics through fieldwork, so a supporting invoice can be linked directly to the sample it supports. Field Agents execute the procedures across the selection; your team reviews the output and owns the judgment. Seniors spend review time on the exceptions and the completeness questions that actually move the opinion, and the mechanical PDF-hunting stays where it belongs, out of the reviewer's day. See how firms run AI agents in audit on the platform, or request a demo to walk through it with your own methodology.