Key Insights
- Tolerable misstatement is the largest error you will accept in one account before concluding it is fairly stated.
- The threshold drives sample size before fieldwork starts: the lower you set it, the more testing the engagement requires.
- Several small errors in the same account can add up past the threshold even when each one looks immaterial on its own.
Set tolerable misstatement too high and you under-test a population that needed scrutiny. Set it without accounting for how small errors add up, and a handful of immaterial misstatements in the same account can quietly push past the threshold and change your conclusion. The number you put on your planning memo decides how much testing you do and how much evidence you need, long before anyone opens a workpaper. This guide covers how tolerable misstatement fits alongside overall materiality and performance materiality, how the threshold drives sample size and evaluation, and how AI changes testing against a fixed number.
What is tolerable misstatement in auditing?
Tolerable misstatement is the cap you set on how wrong a single account can be before you would no longer call it fairly stated. PCAOB materiality guidance treats it as the account- or disclosure-level number you use to assess risk and plan procedures. You set it low enough that the misstatements you find but don't correct, plus the ones you never detect, can't add up to a material misstatement of the financial statements.
Two things set where the number lands. First, it is capped below overall materiality, and below any tighter figure you have set for a specific account. Second, prior-period misstatements inform it: how large they were, what caused them, and what they signal about this year's risk. The threshold is also not locked once set. If facts change mid-engagement and it no longer fits, you reset it and adjust procedures rather than carrying the planning number forward on autopilot.
Sampling makes the number concrete. For substantive tests of details, audit sampling guidance treats tolerable misstatement as the most misstatement you will allow in the population you are testing before concluding it is fairly stated. If your sample covers only part of an account, you hold it to a tighter number than the account's full threshold, because errors in the portion you did not test could combine with the ones the sample turns up.
Tolerable misstatement vs. materiality vs. performance materiality
These three thresholds form a descending hierarchy, from the financial statements as a whole down to the individual account, and each manages aggregation risk at a different level:
- Overall materiality is the statement-level line. Set for the financial statements as a whole, it is the benchmark for planning the nature, timing, and extent of procedures and for judging whether a misstatement could influence a user's decision.
- Performance materiality is the engagement-level buffer. Set below overall materiality, it reduces the chance that uncorrected and undetected misstatements add up past materiality across the audit as a whole.
- Tolerable misstatement is the account-level line. Set below performance materiality and applied at the account, disclosure, or sampling level, it caps how much misstatement a single area can carry before it threatens the statements.
Each threshold guards a different scope. Overall materiality protects the financial statements as a whole, performance materiality guards against misstatements building up across the audit, and tolerable misstatement protects a single account.
Below all three sits a "clearly trivial" floor, defined in misstatement evaluation guidance. Amounts under it are so small they would not matter even if every one of them were added together, so you do not record them at all. Everything above it gets recorded and evaluated.
How tolerable misstatement sets your audit sample size
Tolerable misstatement decides how much testing the engagement needs before anyone touches a workpaper. The lower the threshold, the more assurance the test has to deliver, and the larger the sample required to get there. Say overall materiality on a set of financial statements is $1 million, and you set tolerable misstatement for accounts receivable at $300,000. Drop that to $150,000 because the account carries more risk, and the same population now demands a larger sample, because you are asking the test to catch smaller errors.
Risk of incorrect acceptance and sample size
Sample size depends on more than tolerable misstatement. It also depends on the allowable risk of incorrect acceptance: the risk you are willing to take of concluding a population is fine when its true misstatement actually exceeds tolerable misstatement. Under monetary-unit sampling, that relationship is inverse: the less risk you will accept, the more items you have to test.
The combination is what matters. A tight tolerable misstatement paired with a low acceptable risk produces the biggest samples, and that pairing usually shows up exactly where risk assessment already flagged the account.
Evaluating misstatements against the threshold
At evaluation, the threshold tells you whether the testing worked. Compare known misstatements, projected misstatements from the sample, and the possibility of undetected misstatement against tolerable misstatement. If the total approaches or exceeds it, the response menu opens up: expand testing, ask management to investigate, evaluate whether more misstatements may exist, or confirm management has adjusted the statements.
Sample-size adequacy stays an inspection-sensitive area, because thin samples can lead to insufficient appropriate evidence. The workpaper that holds up shows the sample size tied directly to the objective, the population, the tolerable misstatement, the allowable risk, and the expected error pattern. The sample reads as a planned decision, not a number dropped in after the fact.
Why aggregation is where tolerable misstatement matters most
An item below tolerable misstatement still has to be evaluated with the rest of the population. Every accumulated and projected error has to be evaluated together against the population threshold. Misstatements above the clearly trivial threshold are accumulated, then uncorrected misstatements are evaluated individually and in combination. The evaluation asks whether the population's actual and projected misstatement stays inside the planned limit.
Picture a revenue account with five cutoff errors, each comfortably under tolerable misstatement on its own. Stacked, they can push the account past the planned threshold and force you to reassess the risk of material misstatement for that area. That is the account-level mechanism by which the number actually matters: errors that each clear individually can fail in combination, and the threshold is what catches the combination.
A small misstatement can still matter qualitatively when it is intentional, involves an illegal payment, affects regulatory compliance, or offsets another significant error. Those facts change the evaluation even when the amount looks small in isolation.
Before you evaluate the final effect of uncorrected misstatements, you also reassess whether materiality still holds given the entity's actual results. That reassessment matters most when actual numbers diverge from the figures you used to set planning materiality.
How AI changes audit testing against a fixed threshold
The threshold does not change when AI enters the picture; what changes is how much of the population you can put a procedure against. AI-driven systems can support full-population analytics on selected procedures, not just sample-based testing, which lets auditors run journal-entry testing earlier, surface unusual transactions across large data pools, and process bank statements or contracts faster than manual review allows.
Running a procedure across the full population changes the nature of sampling risk, because you are no longer inferring from a subset. Audit risk does not go away with it. Auditors still have to evaluate the inputs and outputs of the procedure, with human verification built throughout the AI workflow design:
- Data completeness and accuracy
- Procedure design
- Tool reliability
- Threshold appropriateness
- Exception disposition
Those checks are the auditor's to make, and that split is what the operating model runs on: the auditor sets tolerable misstatement and the testing parameters, and the platform executes against them. Fieldguide is built around that split. Field Auditor executes substantive test procedures against selected samples or defined testing populations, flags exceptions, and documents results with a Trace of its work, plus source references where the surface supports them. On the planning side, AI Risk Assessment supports the judgments that inform those thresholds.
None of this lowers the bar. AI-enabled testing still has to meet the same evidence, materiality, and evaluation requirements that have always applied. The amended sampling standard, AS 2315, takes effect for fiscal years beginning on or after December 15, 2026. It keeps the tolerable misstatement framework intact and adds one clarification: items large enough to exceed the threshold on their own are examined individually rather than left to a sample.
See how Fieldguide executes substantive testing against your thresholds
Fieldguide is an end-to-end AI-native platform built for audit and advisory firms, where practitioners and Field Agents work on every engagement: Field Agents execute, and humans review. You set tolerable misstatement, the sample parameters, and the risk-based methodology; the platform keeps planning, testing, evidence, review, and reporting connected in one place. Practitioners retain professional judgment over thresholds, exceptions, and conclusions, while Fieldguide keeps the link between your threshold, your sample, and your evidence visible instead of scattered across spreadsheets and email. Request a demo to see how substantive testing runs against your firm's own thresholds.