Key insights
- Tolerable misstatement, risk of incorrect acceptance, and expected misstatement drive sample size. Population size barely moves it, so "the population grew" is no defense.
- A judgment sample doesn't earn a smaller number. It drops the measured bound on misstatement, and PCAOB findings hit unsupported control reliance and selection, not arithmetic.
- Analytics can scan a whole population, but coverage isn't projection. Only audit sampling lets you conclude on the items you didn't examine.
A reviewer stops on the revenue sample and asks why it's so small. The senior pulls up the sizing table: every input filled in, the arithmetic clean. But the table was never the issue. The sample was built to lean on a control nobody tested, and you can't recalculate your way out of that.
That's the pattern behind most sample-size findings, and it's what this piece works through: what really drives the number, where sizing collapses under PCAOB inspection, and where technology-assisted testing fits in.
What is audit sample size?
Audit sample size is the number of items you pull from a population (an account balance, a class of transactions, a set of controls) to test instead of examining every item. A sample only works if its results project to the whole population, so the number has to be big enough to support that conclusion. Too small and it won't survive review; too large and you've spent hours you didn't need.
What actually drives audit sample size
Sizing a sample comes down to three inputs: how much error you can accept before the account is materially wrong (tolerable misstatement), how confident you need to be that a material error wouldn't slip past the sample (the risk of incorrect acceptance), and how much error you actually expect to find (expected misstatement). Set those three and the arithmetic mostly follows.
When sizing a substantive detail sample, start with tolerable misstatement and the required risk of incorrect acceptance. Then estimate what the population is likely to contain, which the sampling standard frames around the expected size and frequency of misstatements. All else equal, a smaller tolerable misstatement or a lower risk of incorrect acceptance pushes the sample up. So does a higher expected misstatement.
The input people reach for first is the one that barely counts: population size. The standard gives it little effect except when the population is very small. So "the population grew" rests the whole defense on the input that moves the number least.
Of those three inputs, the risk of incorrect acceptance isn't one you set on its own. It falls out of your broader risk assessment. When inherent risk is high or controls aren't effective, you can't accept much sampling risk, so the sample grows, especially if no other substantive test covers the same objective. Lean on other work that supports a lower assessment and the sample can come back down. The catch: every one of these inputs is a judgment call, and each has to trace back to work the file actually contains.
Statistical vs. non-statistical sampling: when to use each
Both are acceptable, and the procedures and misstatement evaluation are the same either way. What separates them is measurement. A statistical design puts a number on the risk that your sample misled you; a judgment sample doesn't.
That gap shows up on review. With a statistical design, you state the confidence level the sample supports and the reviewer moves on. With a judgment design, you rebuild the defense from the risk assessment every time the sample is questioned. And a judgment sample doesn't buy you a smaller one: applied properly, it usually comes in the same size as a well-designed statistical sample, or larger. Come in materially smaller and it fails on size.
So the choice comes down to the population. Statistical methods, attributes sampling for controls or monetary unit sampling for balances, earn their keep on larger, higher-risk populations where the conclusion needs a quantified defense. Non-statistical fits smaller, lower-risk populations where the overhead isn't worth it. Either way, the selection has to be representative: every item in the population had a real chance of being picked, through random or genuinely haphazard selection. Testing whatever the client uploaded first isn't a sample.
Why audit samples fail PCAOB inspection
Here's the pattern in the inspection reports: the number checks out, but the inputs behind it don't. That's what makes these findings expensive. A sizing error you just recalculate. An unsupported input means re-scoping the substantive work, writing comment-letter responses, and sometimes reopening a conclusion you thought was closed.
Sizing a sample on control reliance you never tested
Take the 2024 Grant Thornton inspection. Revenue and receivables samples were too small to back the conclusion, and the count wasn't the problem. The samples were built to rely on controls, but the firm never tested the controls over the accuracy and completeness of customer order data. The reliance was assumed, not earned. The PCAOB's September 2024 banking Spotlight shows the same move on loan testing: sample sizes leaning on other procedures the file never backed up.
When the controls don't hold, the substantive sample has to grow, and you can't argue your way around it after the fact. That's a timing problem partners feel. The re-work lands late, the reviewer sees the file only after the risk is documented, and the sample-size conversation happens in the worst possible order. And a sample sized perfectly is still worthless if the selection quietly skipped part of the population.
Coverage isn't sampling
This one's everywhere: test the biggest dollar items, stop when coverage feels like enough, and call the result a sample. But that's not sampling. You have no way to project to everything you left untested, and that's exactly where inspectors land: the untested chunk is too big, the sample inadequate. A staff practice alert flags the same thing when revenue testing pulls only amounts over a threshold, only unpaid receivables, or only certain months. An item with no chance of being selected can't be projected to, so what you tested isn't a sample.
The fix is stratifying. Pull the top-dollar items and examine them all, 100%, as specific items rather than as a sample. Then sample what's left on its own, and project only within that piece. Covering most of the dollars tells you about those dollars, and nothing about the rest.
Missing documentation cuts the same way. If you can't test an item, you don't get to quietly drop it. For substantive tests, the practice alert treats an untested item as a misstatement and projects it to the population with the others. Drop it instead, and you've understated your projected misstatement, and inspectors read that as unfavorable evidence.
Audit sampling with data analytics: what changes, what doesn't
The 2024 amendments on technology-assisted analysis kick in for fiscal years starting on or after December 15, 2025. For electronic populations, you have three options: examine everything, pick specific items, or sample. Only sampling lets you project to what you didn't test.
Filters and targeted scans are real risk work, but they change the mechanics, not the judgment. What they do change is your paper trail: instead of a selection worksheet, the file has to show how you built the query, what it returned, and how you followed up. Treat that query output as a projected conclusion and you've just repeated the coverage mistake with better software.
There's real regulatory uncertainty here too. Test 100% of journal entries with AI and you might still get inspection questions: the coverage could count in your favor, or an inspector could want so much detail on the model that you fall back to manual sampling. No surprise adoption is slow. In the 2025 audit report from CPA.com and the AICPA, 89% of firms said they don't use technology for substantive analytics.
On Fieldguide, risk-based sampling assistance stays inside the firm's own methodology. The practitioner sets the parameters, reviews what comes back, and owns the conclusion; the AI only works the population data you give it. Inside the Agent Workforce, Field Auditor executes the test procedures across the sample and documents the results with citations and exception flags. The mechanical part, the selection and the write-up, takes less of your senior's time, so their hours go to the judgment calls instead.
Run sampling inside the engagement workflow
Sampling belongs in the engagement workflow, not in a spreadsheet the file re-imports later. Fieldguide is an end-to-end AI-native platform purpose-built for audit and advisory, and its financial audit workflow ties sample parameters to evidence and projection workpapers inside the same engagement record. Keeping sampling in one system is where the time comes back: in one UHY case study, the firm reported 20–30% less time on engagements, with some tasks dropping from three hours to 15 minutes. Request a demo to see how sampling runs end to end, from selection to projection.