CFA Level I Exam · Estimation and Hypothesis Testing
Sampling Methods and Sampling Bias for CFA Level I
Updated 7 October 2026 · Fact-checked
Sampling means selecting part of a population to draw conclusions about the whole. Probability sampling (simple random, stratified, cluster) gives every member a known chance of selection. Non-probability sampling (convenience, judgmental) does not. Biases such as data snooping, survivorship and look-ahead distort results. To answer questions, match the description to the method or bias.
Understand Sampling Methods and Sampling Bias
A population is every member of the group you care about. A sample is a subset you actually observe. You use the sample to estimate population features such as the mean return. Sampling saves time and cost, and sometimes the full population cannot be observed at all.
Probability sampling gives each member a known, non-zero chance of being selected. This lets you use statistical theory to judge how reliable the estimate is. Non-probability sampling relies on convenience or the researcher's judgment. It is quicker and cheaper, but the sample may not represent the population.
The main probability methods are:
- Simple random sampling: every member has an equal chance of selection, like drawing numbers from a hat.
- Systematic sampling: pick every kth member from a list after a random start.
- Stratified random sampling: divide the population into subgroups (strata) that share a characteristic, then draw a simple random sample from each stratum, with sample sizes proportional to the stratum's relative size in the population. This ensures each subgroup is represented. Very small strata may round to zero, so you may need to round up or set a minimum allocation. It is common in bond index tracking, where you sample bonds by maturity, rating and sector.
- Cluster sampling: divide the population into clusters, each a mini-version of the population, such as geographic regions. Randomly select some clusters, then sample all members (one-stage) or a random sample of members (two-stage) from the chosen clusters. Clusters are chosen at random and the others are ignored.
The key contrast: in stratified sampling you sample from every stratum. In cluster sampling you sample only from the selected clusters. Strata are internally similar. Clusters should each look like the whole population.
Non-probability methods are convenience sampling (use what is easy to get) and judgmental sampling (the researcher picks members based on expertise). Both risk bias.
Sampling can go wrong in two ways. Sampling error is the natural difference between a sample statistic and the population value. It exists even in a well-designed random sample. Sampling bias is a systematic distortion from how the data is chosen or used. Common biases:
- Data snooping (data mining) bias: testing many models or variables on the same data until something looks significant. The result is often a chance pattern that fails out of sample.
- Sample selection bias: the sample excludes some data systematically because data is not available for them.
- Survivorship bias: the sample includes only entities that survived, such as funds still operating. Failed funds drop out, so average returns look too high.
- Look-ahead bias: using information in a test that was not available at the time of the decision, such as year-end reported earnings applied to a decision made earlier in the year.
- Time-period bias: the results depend on the period chosen, which may be too short or may include unusual conditions.
Key formulas to remember
- Systematic sampling interval
- k = population size ÷ sample size
- Choose a random start between 1 and k, then take every kth member.
- Stratified sample size per stratum
- Stratum sample = total sample × (stratum size ÷ population size)
- Proportional allocation keeps the sample mix equal to the population mix. Very small strata may round to zero, so round up or set a minimum allocation.
- Sampling error
- Sampling error = sample statistic − population parameter
- Arises from chance even with a proper random sample. It is not the same as bias.
- Bias to method matching
- Failed entities missing → survivorship; future information used → look-ahead; repeated testing on same data → data snooping; unusual or short period → time-period
- Learn the single clue word for each bias.
How to solve Sampling Methods and Sampling Bias questions
Most questions give a short description and ask you to name the method or bias, or to judge the effect on results. Use this sequence.
- 1Read the stem and decide whether it is about a sampling method or a bias.
- 2For a method, ask: is every member's selection chance known? If not, it is non-probability (convenience or judgmental).
- 3If probability, ask whether the population was split into groups. If every group is sampled, it is stratified. If only some randomly chosen groups are sampled, it is cluster.
- 4If members are taken at fixed intervals from a list, it is systematic. If each member is equally likely with no grouping, it is simple random.
- 5For a bias, look for the clue: missing failed firms (survivorship), information not yet available (look-ahead), many tests on one data set (data snooping), unusual or short window (time-period).
- 6Decide the direction of the effect. Survivorship usually overstates average returns and understates risk.
- 7Eliminate the two wrong options and choose the one that matches the clue.
Quickest way: Clue-word matching
When to use it: Use for any definition-style or scenario-style question where you must name a method or bias.
- Underline the action in the stem: random draw, groups, intervals, convenience, repeated testing.
- Match: all groups sampled = stratified; some groups chosen = cluster; every kth = systematic; easy access = convenience; expert choice = judgmental.
- For biases, match: dead funds gone = survivorship; future data = look-ahead; many tries = data snooping.
- Pick the option that fits and move on in about 90 seconds.
Common mistakes in Sampling Methods and Sampling Bias
Confusing stratified and cluster sampling.
Both split the population into groups, so they sound alike.
Fix: Ask whether every group is sampled. Stratified samples every stratum. Cluster samples only the randomly chosen clusters.
Calling sampling error a bias.
Both make the sample result differ from the truth.
Fix: Sampling error is random chance and shrinks with larger samples. Bias is systematic and does not go away by increasing sample size.
Saying survivorship bias understates returns.
Students forget that the failures are the ones removed.
Fix: Removing poor performers pushes average returns up, so reported performance is overstated.
Mixing up look-ahead bias and data snooping.
Both involve improper use of data in backtests.
Fix: Look-ahead uses information not available at the decision date. Data snooping repeatedly searches one data set until a pattern appears.
Treating convenience sampling as probability sampling.
It seems random because no one is deliberately chosen.
Fix: If selection chances are unknown and based on ease of access, it is non-probability.
Worked examples
Example 1
A researcher wants to estimate the average return of a bond universe. She divides the bonds into groups by rating and maturity, then randomly draws bonds from every group in proportion to the group's share of the universe. Which sampling method is this? A. Cluster sampling B. Stratified random sampling C. Judgmental sampling
Show the solution
- The population is divided into subgroups based on shared characteristics (rating and maturity).
- Bonds are drawn randomly from every subgroup, in proportion to its size.
- Sampling from every subgroup defines stratified sampling. Cluster sampling would sample only some groups.
- Selection is random, so it is not judgmental.
Answer: B. Stratified random sampling
Example 2
An analyst studies the average return of equity funds over the past 15 years using only funds that still exist today. Funds that closed after poor results are not in the database. Which statement is most accurate? A. The average return is likely overstated because of survivorship bias B. The average return is likely understated because of look-ahead bias C. The average return is unaffected because the sample is random
Show the solution
- The sample includes only funds that survived, which points to survivorship bias.
- Funds that closed usually had poor returns, and they are missing.
- Excluding poor performers raises the measured average return.
- Look-ahead bias involves unavailable information, which is not described. A set of survivors is not a random sample of all funds.
Answer: A. The average return is likely overstated because of survivorship bias
Exam tips
- Memorize one clue word per method and bias. Most questions are direct matches.
- For stratified versus cluster, the deciding test is whether all groups are sampled.
- Survivorship bias nearly always makes performance look better and risk look lower.
- Remember that more data reduces sampling error but does not cure bias.
- With no penalty for wrong answers, always answer. If the stem states that selection chances are known, rule out convenience and judgmental sampling, then use the group-sampling test to separate stratified, cluster, systematic and simple random.
Practice questions from Estimation and Hypothesis Testing
- Holding the sample and the test statistic constant, a researcher lowers the significance level of a test from 5% to 1%. The change will most…
- A jackknife procedure is applied to a sample of 20 observations to estimate the variability of a statistic. The number of resampled data set…
- An analyst wants to raise the confidence level of an interval for a population mean from 90% to 99% without changing the sample. The interva…
- An analyst wants to estimate the standard error of the sample median of monthly fund returns, for which no simple closed-form formula is con…
- Compared with a sample of 50 observations, drawing a sample of 200 observations from the same population will most likely cause the standard…
Sampling Methods and Sampling Bias: frequently asked questions
What is the difference between stratified and cluster sampling?
In stratified sampling you divide the population into similar subgroups and sample from every one. In cluster sampling you divide it into groups that each mirror the population, randomly select some clusters, and sample only within those. Stratified aims for representation of each subgroup; cluster aims for lower cost.
What is data snooping bias?
Data snooping, or data mining, happens when you test many variables or models on the same data set until you find a significant result. The pattern may be due to chance, so it often fails on new data. Testing on out-of-sample data helps detect it.
What is survivorship bias?
It occurs when a data set includes only entities that survived, such as funds or firms still in existence. Failed entities are missing, so average returns look higher and risk looks lower than they really were.
Is sampling error the same as sampling bias?
No. Sampling error is the random difference between a sample statistic and the population value, and it falls as sample size grows. Sampling bias is a systematic distortion caused by how data is selected or used, and a bigger sample does not fix it.