AICPA Audit Guide: Sampling Risk, Methods, and Evaluation

Audit sampling standards and methods in the United States are set by AU-C Section 530 for audits of non-issuers and by AS 2315 for audits of issuers, and both frameworks let auditors reach conclusions about a full population by testing a subset, provided the sample is designed, selected, and evaluated in a way that controls sampling risk. The two standards share the same core concepts. The choice of method, statistical or non-statistical, attribute or monetary unit or classical variables, depends on what you are testing and how much precision you need.

Which Standard Applies to Your Engagement

AU-C Section 530, issued by the AICPA’s Auditing Standards Board, governs audits of non-issuers: private companies, nonprofits, and government entities. AS 2315, issued by the Public Company Accounting Oversight Board, governs audits of issuers, meaning publicly traded companies and broker-dealers.1PCAOB. Find Analogous Standards The two cover the same territory but diverge on documentation requirements and regulatory oversight.

If you work on public company engagements, note that the PCAOB has adopted amendments to paragraph .11 of AS 2315, approved by the SEC, with an effective date of December 15, 2026.2PCAOB. AS 2315 Audit Sampling (Effective on 12/15/2026) Most of what follows applies to the AICPA framework under AU-C 530, which is the version most practitioners work with.

Statistical vs. Non-Statistical Sampling

Statistical sampling uses probability theory in both selection and evaluation, so sampling risk can be measured mathematically. Non-statistical sampling, sometimes called judgmental, relies on professional judgment and does not produce a quantified risk measure.3PCAOB. AU Section 530 – Audit Sampling Both are acceptable under AU-C 530. Neither is inherently better. The choice depends on the objective of the test, the nature of the population, and the precision you need in the conclusion.

Sampling Risk

Sampling risk is the chance that a conclusion drawn from a sample would differ from the conclusion you would reach by examining every item in the population. You manage it by setting parameters in planning, not by hoping the sample is representative.

For substantive tests of account balances, sampling risk takes two forms:

  • Risk of incorrect acceptance (Beta risk): concluding a balance is fairly stated when it is materially misstated. This is the dangerous outcome. A material misstatement passes undetected.
  • Risk of incorrect rejection (Alpha risk): concluding a balance is misstated when it is correct. This wastes time but does not produce a flawed opinion.

For tests of controls, the parallel risks are the risk of assessing control risk too low (relying on a control that does not work) and the risk of assessing control risk too high (over-testing an effective control). The first is the serious one because it leads to insufficient substantive testing.

Nonsampling Risk

Nonsampling risk covers every source of error unrelated to sampling itself: using an ineffective procedure, misinterpreting evidence, skipping an item accidentally, or applying the wrong criteria when evaluating results.4eGrove. Sampling Risk vs Nonsampling Risk in the Auditors Logic Process Statistical precision does not compensate for an auditor who misreads an invoice. The controls for nonsampling risk are training, supervision, and detailed review.

Tolerable Misstatement and Tolerable Rate of Deviation

Tolerable misstatement is the maximum monetary error you will accept in an account balance while still concluding it is not materially misstated. It flows from overall materiality and is normally set below planning materiality to leave a cushion for undetected errors elsewhere. It is one of the most influential drivers of sample size in substantive testing: the smaller the tolerable misstatement, the larger the sample needed to detect small errors.3PCAOB. AU Section 530 – Audit Sampling

Tolerable rate of deviation plays the same role for tests of controls. It is the maximum control-failure rate you will accept while still concluding the control operates effectively. A lower tolerable rate demands a larger sample.

Attribute Sampling for Tests of Controls

Attribute sampling estimates how often a control fails. The objective is to decide whether the control is reliable enough to justify planned reliance. If not, substantive testing must expand.

Four planning inputs drive the sample:

  • Population: the complete set of items processed by the control during the period.
  • Control attribute: the specific characteristic being tested, such as whether an invoice bears an authorized signature.
  • Tolerable rate of deviation: the highest failure rate you will accept.
  • Expected population deviation rate: your best estimate of how often the control fails, based on prior-year results or a preliminary sample.

The gap between the tolerable rate and the expected rate is the planned allowance for sampling risk. A narrow gap requires a larger sample. A tolerable rate of 5% with zero expected deviations at 95% confidence needs roughly 65 items from a population over 200; the same confidence level at a 10% tolerable rate drops the requirement to about 35.5HUD Office of Inspector General. Appendix A Attribute Sampling These are starting points, adjusted for engagement-specific factors like first-year audit status or a history of control weaknesses.

After testing, calculate the sample deviation rate and add the allowance for sampling risk to reach the Upper Deviation Limit (UDL). The UDL is the worst-case estimate of the true deviation rate in the population. If the UDL sits at or below the tolerable rate, the control passes. If it exceeds the tolerable rate, reliance on that control must be reduced and additional substantive testing is required.3PCAOB. AU Section 530 – Audit Sampling

Substantive Testing: Monetary Unit Sampling

Where tests of controls ask how often a process fails, substantive tests ask how much money is wrong in an account. Two primary approaches are available under the AICPA framework.

Monetary Unit Sampling (MUS) treats each individual dollar as the sampling unit rather than each physical item. Selection is probability-proportional-to-size, so a $100,000 receivable is 100 times more likely to be picked than a $1,000 receivable. That gives automatic focus on large items without a separate stratification step. MUS works best when misstatements are expected to be few and the test is primarily for overstatement.3PCAOB. AU Section 530 – Audit Sampling

Planning an MUS sample requires three inputs: the assessed risk of incorrect acceptance, tolerable misstatement, and expected misstatement.

When you find a misstatement, projection uses a tainting percentage. Tainting measures how wrong each dollar is within a selected item: divide the misstatement by the item’s book value. An account with a book value of $7,090 overstated by $40 has a tainting percentage of about 0.56%. That percentage is applied to the sampling interval to project likely misstatement across the population. This mechanism makes MUS powerful for overstatement testing but weaker for understatements and useless for zero-balance items, since a $0 book value cannot produce a meaningful tainting percentage.

Substantive Testing: Classical Variables Sampling

Classical Variables Sampling (CVS) uses the physical item as the sampling unit. It is the better choice when misstatements are expected to be numerous, when balances vary widely, or when the test needs to cover both overstatement and understatement. CVS lets you project total monetary misstatement and construct a confidence interval around that estimate.

Three main techniques:

  • Mean-per-unit: estimates the average audited value per item and multiplies by the population size. It requires the largest sample of the three.
  • Difference estimation: calculates the average difference between audited and book values, then projects that difference across the population. Efficient when differences are relatively uniform.
  • Ratio estimation: uses the ratio of total audited to total book values in the sample, then applies that ratio to the full population’s book value. Efficient when misstatements are proportional to book values.

Difference and ratio estimation tend to need smaller samples than mean-per-unit because they use the relationship between audited and recorded amounts to reduce variability. All CVS techniques require an assessment of the population’s standard deviation, and higher variability drives sample size up.

Whatever the method, the risk of incorrect acceptance is set low, commonly at 5% or 10%, because accepting a materially misstated balance is the worst possible audit outcome.3PCAOB. AU Section 530 – Audit Sampling

Defining the Population and Selecting Items

Before pulling a single item, define the sampling unit precisely (a canceled check, an invoice line, an individual dollar) and set the boundaries of the population. AU-C 530 requires evidence that the population is complete and appropriate for the audit objective.6AICPA. Audit Sampling – AU-C Section 530 For a completeness test of revenue, the population should be shipping documents, not recorded sales invoices, because starting from invoices misses shipments that were never billed.

Stratification is one of the most effective design techniques available. Dividing the population into subgroups (for instance, receivables above and below $50,000) concentrates testing where the risk of material misstatement is highest and uses smaller samples for lower-risk strata. MUS builds stratification into its selection through probability-proportional-to-size.

Selection Methods

The selection method must match the sampling approach. For statistical sampling:

  • Random selection: every sampling unit has equal probability of being chosen. Random number generators remove bias. It is the cleanest method for statistical samples.
  • Systematic selection: pick a random starting point, then select every nth item using an interval calculated as population size divided by desired sample size. This works well but produces biased results if the population has a pattern aligned with the interval.

For non-statistical sampling, haphazard selection is acceptable. Items are chosen without a structured technique but with deliberate avoidance of conscious bias. Haphazard selection cannot be used for statistical sampling because it gives no basis for measuring the probability of selection.3PCAOB. AU Section 530 – Audit Sampling

Missing Items and Anomalies

If you select an item and the supporting documentation is missing, AU-C 530 is clear: when you cannot apply the planned procedure and no suitable alternative exists, treat the item as a deviation (for control tests) or a misstatement (for substantive tests).6AICPA. Audit Sampling – AU-C Section 530 You do not skip the item. The standard allows a practical exception: if treating the unexamined item as misstated would not change the overall sample evaluation, further investigation may not be necessary. In most cases where a missing item matters, perform alternative procedures or accept the added misstatement or deviation.

International standards (ISA 530) recognize an “anomaly,” a misstatement demonstrably not representative of the population, and allow excluding it from population projection under strict conditions requiring a high degree of certainty that the error is isolated. The AICPA approach is more conservative. The Auditing Standards Board removed the anomaly concept from AU-C 530, so under U.S. non-issuer standards all sample misstatements are generally projected to the population. Firms vary in their internal policies, but the safe course is to project every misstatement unless extraordinary evidence supports isolation. Even when excluded from projection, the misstatement’s effect must still be considered in the overall evaluation of the financial statements.

Evaluating and Documenting Results

For tests of controls, the evaluation is straightforward: calculate the sample deviation rate, add the allowance for sampling risk to reach the UDL, and compare the UDL to the tolerable rate. At or below, the control passes. Above, the audit plan changes.

For substantive tests, project the misstatement found in the sample to the whole population. MUS uses tainting percentages. CVS uses the mean, difference, or ratio estimate to produce a point estimate and confidence interval. The projected misstatement plus an allowance for sampling risk is compared to tolerable misstatement. Below it, the balance is fairly stated. Above it, the account is likely materially misstated.3PCAOB. AU Section 530 – Audit Sampling

When projected misstatement exceeds tolerable misstatement, options include asking management to investigate and correct the identified errors, expanding the sample to narrow the allowance for sampling risk, or performing different substantive procedures targeting the same assertion.

Qualitative Evaluation

Numbers do not tell the whole story. Evaluate the nature and cause of every misstatement, regardless of dollar amount. A $500 transposition and a $500 deliberate document alteration are not the same finding. An intentional misstatement can be material for qualitative reasons even when small, because it may signal management bias or fraud risk.7PCAOB. Auditing Standard 14 Appendix B – Qualitative Factors Related to the Evaluation of the Materiality of Uncorrected Misstatements Consider whether misstatements point to contract violations, conflicts of interest, or an unwillingness to fix known weaknesses in financial reporting.

Documentation

AU-C 530 requires comprehensive documentation across planning, execution, and evaluation. The most common sampling-related deficiency in peer reviews is failure to justify or determine sample size adequately, followed by failure to link the testing back to the risk assessment.

At a minimum, cover:

  • Test objective and population definition: what is being tested and what constitutes the full population.
  • Sampling method and size determination: statistical or non-statistical, the inputs used (tolerable misstatement, risk of incorrect acceptance, expected error rate), and how they produced the sample size.
  • Selection details: how items were chosen, which items were examined, and the nature and cause of any deviations or misstatements found.
  • Projection and conclusion: the calculated projected misstatement or Upper Deviation Limit, and the explicit conclusion about the population.3PCAOB. AU Section 530 – Audit Sampling