Statistical sampling in auditing is the practice of testing a randomly selected subset of items from an account balance or class of transactions and using probability theory to draw a measurable conclusion about the full population. It lets an auditor say, with a defined level of confidence, how much misstatement or control failure is likely to exist without examining every item. Two methods dominate: attribute sampling for tests of controls, and variable sampling (usually monetary unit sampling or classical variables sampling) for substantive tests of dollar balances.
What Makes a Sample Statistical
Statistical sampling means applying an audit procedure to less than 100 percent of the items in an account balance or class of transactions, where every item has a known chance of being selected.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling Two requirements separate it from non-statistical sampling: the selection must be random, and the results must be evaluated using statistical methods. Miss either one and the sample is non-statistical regardless of how large it is.
Non-statistical (or judgmental) sampling relies on the auditor’s experience to pick items. It has its place, but it cannot produce a number representing the probability that the conclusion is wrong. Statistical sampling can. An auditor using statistical methods can state that there is no more than a 5 percent chance the sample isn’t representative of the population, and that statement is mathematically grounded rather than intuitive.2Office of the Comptroller of the Currency. Comptrollers Handbook – Sampling Methodologies
The practical payoff is efficiency with disclosed uncertainty. Instead of examining tens of thousands of invoices, the auditor tests a scientifically determined number and reaches a conclusion about the entire account, carrying a known margin of error that is managed rather than ignored.
The Concepts That Drive Every Sampling Decision
A handful of ideas shape every choice inside a sampling plan. Get these right and the mechanics follow.
Population
The population is the complete set of items the auditor wants to draw a conclusion about: every sales invoice for the year, every inventory tag in a warehouse, every disbursement from a given account. The auditor has to verify the population is complete before pulling a sample. Testing from an incomplete population produces conclusions about the wrong universe of transactions.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling
Sampling Risk
Sampling risk is the possibility that the auditor’s conclusion based on the sample differs from the conclusion that would be reached by testing every item.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling Larger samples reduce it. Only a full examination eliminates it.
In substantive testing, sampling risk takes two forms. The risk of incorrect acceptance is the danger of concluding an account balance is fairly stated when it actually contains a material misstatement. The risk of incorrect rejection is the opposite. Incorrect acceptance is the far more consequential error: incorrect rejection just leads to extra work, while incorrect acceptance sends a flawed audit opinion out the door. For tests of controls, the parallel risks are assessing control risk too low (relying on a control that isn’t working) and assessing control risk too high (distrusting a working control, which triggers unnecessary additional testing).1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling
Non-Sampling Risk
Non-sampling risk is every part of audit risk that has nothing to do with sample size. An auditor can examine every item in a population and still miss a material misstatement if the wrong procedure was used or a problem in the documents wasn’t recognized. Confirming recorded receivables, for example, won’t reveal receivables that were never recorded. That’s a design failure, not a sample-size failure. Planning, supervision, and quality control reduce non-sampling risk to a negligible level.3Public Company Accounting Oversight Board. AU Section 350.11 – Audit Sampling
Confidence Level
Confidence level is the flip side of sampling risk. A 95 percent confidence level means the auditor accepts no more than a 5 percent risk that the sample isn’t representative.2Office of the Comptroller of the Currency. Comptrollers Handbook – Sampling Methodologies Higher confidence requires a larger sample. For tests of controls where the auditor plans heavy reliance, 90 or 95 percent is typical.4U.S. Department of Housing and Urban Development Office of Inspector General. Appendix A – Attribute Sampling For a substantive test, a lower initial confidence level may be acceptable when other procedures such as analytical review address the same assertion.
Tolerable Misstatement and Tolerable Deviation Rate
In substantive testing, tolerable misstatement is the largest dollar error that can exist in an account without making the financial statements materially misstated. It flows from the materiality judgments set in planning. In tests of controls, the equivalent is the tolerable deviation rate: the highest failure rate the auditor can accept while still concluding the control is reliable. Five percent is a common threshold, though it varies with how important the control is.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling
How Auditors Set Sample Size
Sample size isn’t guessed. It comes from the interplay of several inputs, and getting it wrong in either direction either wastes hours or manufactures false assurance.
For substantive tests, four factors drive the number: the tolerable misstatement, the allowable risk of incorrect acceptance (which reflects inherent risk, control risk, and any other procedures covering the same assertion), the expected size and frequency of misstatements, and the characteristics of the population. The relationships are intuitive: a smaller tolerable misstatement, a higher assessed risk, or a population with expected large or frequent errors all push the sample up. For tests of controls, the parallel inputs are the tolerable deviation rate, the likely deviation rate, and the allowable risk of assessing control risk too low.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling
In practice, auditors read from published sample-size tables or use software. As an illustration, attribute testing at 95 percent confidence with a 5 percent tolerable deviation rate and zero expected deviations calls for a minimum sample of about 65 items. Drop the confidence to 90 percent and it falls to about 50. Widen the tolerable rate to 10 percent at 90 percent confidence and only about 25 items are needed. Those figures assume a population above 200 items.4U.S. Department of Housing and Urban Development Office of Inspector General. Appendix A – Attribute Sampling
Stratification
Stratification is one effective way to shrink a sample without giving up assurance. The auditor splits the population into more uniform subgroups based on characteristics such as recorded dollar value, then samples each subgroup separately.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling Individually significant items (those exceeding tolerable misstatement on their own) are pulled out and examined 100 percent, leaving a lower-variability group to sample from.
Attribute Sampling for Tests of Controls
Attribute sampling answers a single question: how often does a control fail? The auditor defines a specific attribute (a required signature, a matching purchase order number, evidence of a credit check) and counts how many sample items lack it. The result is a deviation rate rather than a dollar amount.
The steps run in a set order. Set the confidence level and tolerable deviation rate. Estimate the expected deviation rate. Read the sample size from a table. Select items randomly. Inspect each one. Compute the upper deviation rate, which is the highest likely failure rate in the full population given what the sample showed.
If the upper deviation rate falls below the tolerable deviation rate, the auditor concludes the control is reliable enough to lean on. If it exceeds the tolerable rate, the control can’t be relied on and substantive testing for the related assertion has to expand.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling
Variable Sampling for Substantive Tests
Where attribute sampling measures rates, variable sampling measures dollars. The purpose is deciding whether an account balance is materially misstated. Two techniques dominate.
Monetary Unit Sampling
Monetary unit sampling (MUS), also called probability-proportional-to-size sampling, is the most widely used statistical method for substantive testing. The sampling unit is the individual dollar. Every dollar has an equal chance of selection, so larger transactions are proportionally more likely to be picked. A $100,000 receivable is ten times more likely to appear in the sample than a $10,000 receivable.
That built-in stratification is the practical advantage. High-value items get the heaviest scrutiny without the auditor having to create strata by hand. MUS also doesn’t require estimating the population’s standard deviation up front, which keeps planning simpler.
The trade-off is a blind spot. Because low-value items have low selection probabilities, MUS catches overstatements better than understatements. An item recorded at $500 that should be $50,000 is unlikely to be picked up. MUS therefore fits best on asset and revenue accounts where the primary risk is overstatement.
Classical Variables Sampling
Classical variables sampling (CVS) selects physical items (invoices, balances, journal entries) rather than individual dollars. Each record has the same probability of selection regardless of its recorded size. CVS uses normal distribution theory to estimate the total population value or total misstatement, through techniques such as mean-per-unit, difference, or ratio estimation.
CVS handles understatements without the structural weakness of MUS since selection doesn’t depend on recorded value. It’s the better fit when misstatements are expected across all value ranges or when understatement is the main risk. The cost is complexity: the auditor must estimate the population’s standard deviation before sizing the sample, and the evaluation math is heavier.
How the Sample Is Selected
Statistical sampling requires genuine randomness. The auditor can’t grab items that look interesting or convenient. Two methods dominate.
With random number selection, each item in the population gets a unique number and a random number generator picks which items to test. Human bias is removed entirely. Every item has an independently determined chance of selection, and selecting one item doesn’t change the odds for another.
With systematic selection using a random start, the auditor divides the total population (in units or dollars) by the sample size to get a sampling interval. A random number within the first interval becomes the starting point, and every subsequent item at that interval distance enters the sample. For MUS this walks through the cumulative dollar total: with an interval of $50,000 and a random start of $12,000, the selected dollar positions are $12,000, $62,000, $112,000, and so on, and the transactions containing those positions get pulled. Systematic selection is efficient, but the auditor should confirm the population isn’t ordered in a cycle that lines up with the interval, because a matching cyclical pattern can produce a badly skewed sample.
Haphazard selection, where an auditor picks items without conscious bias but without a formal random mechanism, is a non-statistical technique. Research has shown that haphazard samples consistently differ from truly random ones because subconscious tendencies over- and under-select certain items. Auditing standards allow haphazard selection for non-statistical sampling, but it can’t support the probability-based conclusions that statistical sampling requires.
Evaluating and Projecting the Results
The final step is projecting sample findings back to the full population. A few errors in a sample don’t mean the account is misstated by exactly that amount. The findings have to be extrapolated.
For MUS, the auditor calculates a projected misstatement by extrapolating each error to the sampling interval that produced it. The tainting percentage (the misstatement divided by the recorded value of the sampled item) is multiplied by the interval to estimate how much misstatement likely sits in that portion of the population. The projected misstatement is then combined with an allowance for sampling risk to produce the upper misstatement limit, which represents the maximum misstatement likely to exist in the account at the chosen confidence level. Compare that limit to the tolerable misstatement set during planning. At or below, the account is considered fairly stated. Above, it’s considered materially misstated.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling
There is a gray zone. When the projected misstatement is close to tolerable misstatement but doesn’t clearly exceed it, the auditor may still conclude the risk that actual misstatements exceed the tolerable amount is unacceptably high.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling Near-misses call for skepticism, not relief.
For attribute sampling, the auditor counts the deviations and computes the upper deviation rate from tables or software, then compares it to the tolerable deviation rate. Two or more deviations in a test designed for a 5 percent tolerable rate can be enough for the auditor to conclude the true deviation rate is unacceptably likely to exceed 5 percent.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling
What Happens When Results Exceed Tolerable Thresholds
Unfavorable results don’t end the audit. They redirect it.
When a test of controls produces a deviation rate above the tolerable threshold, the auditor has to reassess the planned reliance on that control and expand substantive testing for the affected assertions.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling In practice that means more detailed transaction-level work to compensate.
When a substantive test produces a projected misstatement higher than expected, the auditor should reconsider the risk assessments that shaped the plan. Higher-than-anticipated errors in one area often signal problems elsewhere, and tests designed on the same risk assumptions may need to be adjusted.1Public Company Accounting Oversight Board. AS 2315 – Audit Sampling Other responses include increasing the sample to narrow the allowance for sampling risk, asking management to investigate and correct the identified misstatements, and, if material misstatements remain uncorrected, modifying the audit opinion.
The discipline of statistical sampling shows up here. Because the auditor started with defined thresholds and measurable risk levels, the response follows a logical path. The numbers say what happened; the standards say what to do about it.