Skip to content

Level III Core · Portfolio Performance Evaluation

Benchmarks and Their Selection for CFA Level III

Updated 9 October 2026 · Fact-checked

A benchmark is a standard against which you judge a portfolio's return and risk. A valid benchmark is unambiguous, investable, measurable, appropriate, specified in advance, accountable and reflective of current investment opinions. To solve questions, test each property against the mandate, then judge benchmark quality using the quality tests.

Understand Benchmarks and Their Selection

A benchmark is a reference portfolio that represents the manager's investment mandate. You compare the portfolio with it to see whether the manager added value. Without a good benchmark, any return number is hard to interpret.

The CFA curriculum lists the properties of a valid benchmark:

  • Specified in advance: set at the start of the evaluation period, not chosen later.
  • Appropriate: consistent with the manager's investment style or area of expertise.
  • Measurable: its value can be calculated frequently, at least as often as the portfolio is evaluated.
  • Unambiguous: the names and weights of the securities are clearly identified.
  • Reflective of current investment opinions: the manager has opinions on the securities in it.
  • Accountable: the manager accepts the benchmark as a fair representation of the mandate.
  • Investable: it is possible to replicate the benchmark by holding its securities.

The main types of benchmarks are absolute return benchmarks (a fixed target such as a required return), manager universes (median manager of a peer group), broad market indexes, style indexes, factor-model-based benchmarks, returns-based benchmarks, and custom security-based benchmarks. Each type has strengths and weaknesses. Broad market indexes are generally unambiguous, investable and easy to obtain, but they may not match the manager's style. Manager universes are measurable but are not unambiguous, not investable and not specified in advance, and they suffer from survivorship bias and classification bias. Custom benchmarks fit the mandate best but cost more to build.

You assess benchmark quality with diagnostics on how well it fits the manager's actual portfolio. The standard tests are:

  • Systematic bias: regress the portfolio's return on the benchmark's return. The benchmark is unbiased if the beta is close to 1 and the alpha is close to 0. You can run the same test on active return instead. Regress the active return on the benchmark's return, and the slope should be close to 0 (it equals the beta minus 1). Equivalently, the correlation between active return and benchmark return should be low. A slope clearly away from 0 (or a beta clearly away from 1) means the benchmark has a systematic bias relative to the manager's style.
  • Tracking error: the standard deviation of active return. It should be consistent with the manager's intended active risk. If it is much higher, the cause may be a poor benchmark fit or the manager taking more active risk than intended, so you need other evidence to tell which.
  • Style fit (risk characteristic match): the portfolio's characteristics, such as valuation, size and risk exposures, should match the benchmark's. A persistent gap indicates poor benchmark fit or a style mismatch. A gap that widens over successive periods indicates style drift in the manager's portfolio.
  • Benchmark turnover: it should be low, because high turnover makes the benchmark costly to replicate.
  • Coverage ratio: the share of the portfolio's market value that is in securities held in the benchmark. A high ratio indicates good fit.

Active share is not a test of benchmark quality. It is a measure of manager activity: how much the portfolio's holdings deviate from the benchmark weights.

A custom security-based benchmark is built from the securities the manager actually covers, with weights set to reflect the mandate. Steps: define the investment universe, apply the manager's screens (such as size, liquidity or style), choose a weighting scheme (market cap, equal weight or other), and rebalance at a set frequency. The result is tailored and transparent but needs ongoing maintenance and is costly.

Key rules to remember

Active return
Active return = Rp − Rb
Portfolio return minus benchmark return over the same period.
Tracking error
Tracking error = standard deviation of (Rp − Rb)
Measures how closely the portfolio follows the benchmark. Tracking error well above the manager's intended active risk may reflect a poor benchmark fit or extra active risk taken by the manager.
Systematic bias test
Rp = a + b × Rb + error, or equivalently (Rp − Rb) = a + (b − 1) × Rb + error
Regress portfolio return on benchmark return: a good benchmark has b close to 1 and a close to 0. In the active return form, the slope (b − 1) should be close to 0. A slope clearly away from 0 signals benchmark bias.
Information ratio
IR = average active return ÷ tracking error
Uses the benchmark as the reference, so a poor benchmark distorts it.
Benchmark properties checklist
Specified in advance, Appropriate, Measurable, Unambiguous, Reflective of current opinions, Accountable, Investable (SAMURAI)
A memory aid for the seven properties of a valid benchmark.

How to solve Benchmarks and Their Selection questions

Use this method for any question on benchmark selection or quality. Always anchor your answer to the mandate in the vignette.

  1. 1Read the mandate: asset class, style, constraints, and what the client wants the benchmark to measure.
  2. 2List the candidate benchmarks and identify each type (absolute, broad index, style index, universe, custom, returns-based).
  3. 3Test each candidate against the seven properties. Name the property that fails.
  4. 4Check quality diagnostics if data are given: systematic bias (beta of portfolio return on benchmark return near 1, or slope of active return on benchmark return near zero), tracking error relative to the manager's active risk, style fit, benchmark turnover and coverage ratio.
  5. 5Decide whether to keep, adjust or build a custom benchmark, and say why in one sentence tied to the mandate.
  6. 6If asked to construct a custom benchmark, give the steps: universe, screens, weighting, rebalancing, and reconstitution rules.
  7. 7Answer exactly the command word: identify, state, justify or recommend. Give only the number of items asked for.

Quickest way: Property-failure scan

When to use it: Use when an item-set question asks which benchmark is most appropriate or which property is violated.

  1. Underline the manager's style and universe in the vignette.
  2. Eliminate any benchmark that does not match that style or universe (fails appropriate).
  3. Eliminate any that cannot be held or priced regularly (fails investable or measurable).
  4. Eliminate any chosen after the fact or with unclear holdings (fails specified in advance or unambiguous).
  5. Pick the one remaining. If two remain, choose the one with the closer style fit.

Common mistakes in Benchmarks and Their Selection

  • Treating a manager universe as a valid benchmark.

    Peer comparison feels natural and is common in practice.

    Fix: Remember that universes are not unambiguous, not investable and not specified in advance, and they suffer from survivorship and classification bias.

  • Using a broad market index for a style-specific manager.

    Broad indexes are easy to find and widely quoted.

    Fix: Match the benchmark to the style. A small-cap value manager measured against a large-cap index shows misleading active return.

  • Confusing 'investable' with 'measurable'.

    Both sound like practical features.

    Fix: Investable means you could hold the securities to replicate it. Measurable means its return can be calculated at the evaluation frequency.

  • Forgetting that the benchmark must be specified in advance.

    Candidates focus on fit and overlook timing.

    Fix: A benchmark picked after seeing results fails this property, even if it fits well.

  • Listing properties without linking them to the case.

    Candidates memorise the list and stop there.

    Fix: Name the failing property and quote the fact from the vignette that causes the failure.

  • Writing more items than the question asks for.

    Candidates hope extra answers will earn credit.

    Fix: Only the requested number of responses is evaluated, in the order given. Give exactly that number.

Worked examples

Example 1

A manager invests only in mid-cap value stocks. The client proposes a broad large-cap index as the benchmark. Identify the property the benchmark most clearly violates and recommend an alternative.

Show the solution
  1. The manager's style is mid-cap value, so the benchmark should reflect that.
  2. A large-cap index holds different securities and a different style, so it is not appropriate for the mandate.
  3. It also would not reflect the manager's current investment opinions, because the manager does not follow most of its stocks.
  4. An alternative is a mid-cap value style index, or a custom benchmark built from the manager's mid-cap value universe.

Answer: The benchmark is not appropriate for the manager's style. Use a mid-cap value index or a custom benchmark built from the manager's universe.

Example 2

Over several years, a regression of a manager's active return on the benchmark's return gives a slope of 0.45. The annualised tracking error is 9%, while the manager's stated active risk target is 3%. Assess the benchmark's quality.

Show the solution
  1. For a good benchmark, the slope of active return on benchmark return should be close to zero. A slope of 0.45 is clearly positive. It is the same as a beta of 1.45 for portfolio return on benchmark return (0.45 + 1 = 1.45), well above 1.
  2. A positive slope means the portfolio's active return rises when the benchmark rises. The benchmark has a systematic bias relative to the manager's style.
  3. Tracking error of 9% is three times the 3% target (9 ÷ 3 = 3), so the portfolio does not follow the benchmark as closely as intended.
  4. High tracking error alone is ambiguous. It may reflect a poor benchmark fit, or the manager taking more active risk than intended. Style fit and coverage ratio would help tell which.
  5. The slope is direct evidence of bias in the benchmark. Active return measured against this benchmark will mix manager skill with a style mismatch.

Answer: The slope of 0.45 shows systematic bias, so the benchmark is a poor fit. Tracking error at three times the target supports this concern, but on its own it may reflect either poor fit or extra active risk by the manager. Replace or adjust the benchmark, for example with a custom benchmark that matches the manager's actual exposures.

Exam tips

  • Memorise the seven properties and be ready to name the exact one that fails in a vignette.
  • When asked to justify, link the property to a specific fact in the case. One sentence is enough.
  • For custom benchmark construction, answer in the order: universe, screens, weighting, rebalancing.
  • Do not claim a universe or peer median is investable. That is a frequent trap.
  • Know that the systematic bias test regresses portfolio return on benchmark return (beta near 1, alpha near 0), or active return on benchmark return (slope near 0, low correlation). High tracking error may mean poor fit or extra active risk by the manager, so look for other evidence.
  • If the question asks for two properties, give exactly two. Extra answers are not marked.

Benchmarks and Their Selection in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Benchmarks and Their Selection: frequently asked questions

What are the qualities of a valid benchmark in CFA Level III?

A valid benchmark is specified in advance, appropriate, measurable, unambiguous, reflective of current investment opinions, accountable and investable. Questions usually ask you to spot which one is violated in a case.

How do I choose an appropriate benchmark for a portfolio?

Start with the manager's mandate and style. Choose the benchmark that matches that style and can be held and measured regularly. If no index fits, build a custom benchmark from the manager's universe.

What are the main types of benchmarks?

They include absolute return benchmarks, manager universes, broad market indexes, style indexes, factor-model-based benchmarks, returns-based benchmarks and custom security-based benchmarks. Know one strength and one weakness of each.

How do I test benchmark quality?

Check for systematic bias by regressing portfolio return on benchmark return (beta near 1) or active return on benchmark return (slope near zero). Then check that tracking error is consistent with the manager's intended active risk, style characteristics match, benchmark turnover is low and coverage ratio is high. Poor results suggest the benchmark does not fit the mandate.