CFA Level I Exam · Introduction to Financial Data Science
Data Types and Data Preparation for CFA Level I
Updated 7 October 2026 · Fact-checked
Data types describe how data is organized: structured (tables with fixed fields), unstructured (text, images, audio) and semi-structured (tagged but flexible, such as JSON). Data preparation cleans errors, then transforms variables. Normalization rescales to 0-1 using min and max. Standardization rescales to mean 0 and standard deviation 1.
Understand Data Types and Data Preparation
Data science starts with data, and data comes in different shapes. Structured data fits in rows and columns with defined fields, like a table of daily closing prices, or a spreadsheet of company ratios. Unstructured data has no predefined format: news articles, earnings call transcripts, social media posts, images, audio and satellite photos. Semi-structured data sits between the two. It has some tags or markers that give organization, but the layout is flexible. JSON and XML files, and HTML web pages, are common examples.
Models need clean, consistent input. Raw data is rarely ready. Data preparation (also called data cleansing and preprocessing) turns raw data into usable data. Cleansing deals with errors. Typical problems are missing values, invalid values, inaccurate values, non-uniform formats (for example dates written in different styles or currencies mixed together) and duplicate observations. Different problems have different fixes. You may delete the observation, replace the value with an estimate such as the mean or median, or correct the value after checking the source.
Outliers are extreme values far from the rest. First decide whether they are errors or real. You can find them with a standard-deviation rule (z-score) or the interquartile range (IQR) rule. You can then remove them, replace them with a boundary value (winsorization), or trim them (drop the tails). A real extreme return may carry information, so do not delete it blindly.
After cleansing, you transform the variables. Normalization (min-max scaling) squeezes values into the range 0 to 1. It is sensitive to outliers because the min and max set the scale. Standardization centers values on a mean of 0 and a standard deviation of 1. It is less affected by outliers and is useful when data is roughly normal. Scaling matters because variables with large units, such as market cap in billions, can swamp variables with small units, such as a ratio of 0.5.
For text, preparation is different. You clean by removing punctuation, numbers and stop words, then split text into tokens and reduce words to their stems or lemmas. That belongs to text analytics, but the same idea holds: turn messy raw input into structured form.
Key formulas to remember
- Normalization (min-max scaling)
- X_norm = (X − X_min) ÷ (X_max − X_min)
- Result lies between 0 and 1. Sensitive to outliers because min and max are used.
- Standardization (z-score scaling)
- X_std = (X − μ) ÷ σ
- Result has mean 0 and standard deviation 1. It does not bound values to a fixed range.
- IQR outlier rule
- Outlier if X < Q1 − 1.5 × IQR or X > Q3 + 1.5 × IQR, where IQR = Q3 − Q1
- Common rule of thumb; a larger multiplier such as 3 flags only extreme outliers.
- Cleansing problems checklist
- Missing, invalid, inaccurate, non-uniform, duplicate
- Know the example for each type and its usual fix.
How to solve Data Types and Data Preparation questions
Use this method for any question on data types or data preparation.
- 1Read the stem and identify what is being asked: classify data, name a cleansing problem, or choose a transformation.
- 2For classification, ask: does it have fixed fields in rows and columns (structured), no set format (unstructured), or tags with flexible layout (semi-structured)?
- 3For a cleansing problem, match the symptom to the label: blank entries are missing; impossible entries (a negative share count) are invalid; wrong but plausible values are inaccurate; mixed formats are non-uniform; repeated records are duplicates.
- 4For outliers, decide whether the value is an error or genuine, then pick removal, trimming or winsorization.
- 5For scaling, check whether the answer needs a 0-1 range (normalization) or mean 0 and standard deviation 1 (standardization).
- 6If a calculation is needed, plug into the formula and check that the answer is in the expected range.
- 7Eliminate the two weaker options by checking each against the definition.
Quickest way: Keyword matching for data questions
When to use it: Use this for definition and classification questions, which you should answer in well under 90 seconds.
- Spot the clue word: table or database means structured; text, image or audio means unstructured; JSON, XML or tags means semi-structured.
- For scaling, 0-to-1 means normalization; mean 0 and unit variance means standardization.
- For calculations, normalization needs min and max; standardization needs mean and standard deviation.
- Sanity check: a normalized value can never be below 0 or above 1 when computed on the same data.
Common mistakes in Data Types and Data Preparation
Calling JSON or XML data unstructured.
It looks like text, so it feels unstructured.
Fix: If it has tags or key-value pairs that give organization, it is semi-structured.
Mixing up normalization and standardization.
Both are called scaling and the names sound similar.
Fix: Normalization uses min and max and gives 0 to 1. Standardization uses mean and standard deviation and gives mean 0 and standard deviation 1.
Deleting every outlier automatically.
Outliers look like errors.
Fix: Check first. A genuine extreme return may be informative. Remove or adjust only if it is an error or distorts the purpose of the analysis.
Confusing invalid values with inaccurate values.
Both mean the value is wrong.
Fix: Invalid values fall outside a meaningful range (a negative age). Inaccurate values are plausible but wrong (a wrong price that still looks reasonable).
Forgetting that normalization is sensitive to outliers.
Students remember the formula but not the effect.
Fix: One extreme maximum stretches the range and compresses all other values toward 0. Standardization is less affected.
Dividing by the wrong quantity in the min-max formula.
Students divide by the maximum instead of the range.
Fix: The denominator is X_max − X_min.
Worked examples
Example 1
A dataset of a company's P/E ratios has a minimum of 8, a maximum of 28, a mean of 16 and a standard deviation of 5. Using min-max scaling, the normalized value of a P/E of 12 is closest to: A) 0.20, B) 0.40, C) 0.80
Show the solution
- Use X_norm = (X − X_min) ÷ (X_max − X_min).
- Numerator: 12 − 8 = 4.
- Denominator: 28 − 8 = 20.
- Result: 4 ÷ 20 = 0.20.
- Option B (0.40) comes from 8 ÷ 20, which uses the minimum as the numerator. Option C (0.80) comes from 16 ÷ 20, which uses the mean as the numerator. Only 4 ÷ 20 follows the formula.
Answer: A) 0.20
Example 2
Using the same dataset (mean 16, standard deviation 5), the standardized value of a P/E of 26 is closest to: A) 0.5, B) 2.0, C) 10.0
Show the solution
- Use X_std = (X − μ) ÷ σ.
- Numerator: 26 − 16 = 10.
- Divide by 5: 10 ÷ 5 = 2.0.
- Option A (0.5) comes from 5 ÷ 10, which inverts the division. Option C (10.0) comes from 26 − 16 and stops there, without dividing by the standard deviation.
Answer: B) 2.0
Exam tips
- Expect definition and classification questions. Memorize one example for each data type and each cleansing problem.
- Know the contrast: normalization is bounded 0 to 1 and outlier-sensitive; standardization is not bounded and handles outliers better.
- In a calculation, do the subtraction first and write it down. Mistakes usually come from wrong inputs, not from division.
- Read for the data's purpose. A question may ask which step comes before scaling: cleansing comes first.
- With no penalty for wrong answers, always answer. Eliminate options that mismatch the definition, then choose.
Practice questions from Introduction to Financial Data Science
- A researcher wants to show the distribution, median, quartiles and outliers of monthly returns for several funds side by side. The visualiza…
- A model predicting loan defaults achieves very low error on the training data but performs poorly on new data. This outcome is best describe…
- A data scientist wants to see how the term frequency of the words appearing in a set of central bank statements compares across words, using…
- An asset manager wants to analyze thousands of earnings call recordings and news articles alongside price data. Which characteristic of big …
- A feature has values ranging from 10 to 50, with a minimum of 10 and a maximum of 50. An analyst applies min-max normalization. The normaliz…
Data Types and Data Preparation in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Data Types and Data Preparation: frequently asked questions
What is the difference between structured and unstructured data?
Structured data fits in a fixed table with defined fields, such as price histories or financial statement line items. Unstructured data has no predefined format, such as text, images and audio. Semi-structured data, such as JSON, has tags but a flexible layout.
What is the difference between normalization and standardization?
Normalization rescales values to a 0-to-1 range using the minimum and maximum. Standardization rescales values to have a mean of 0 and a standard deviation of 1 using the mean and standard deviation. Normalization is more sensitive to outliers.
How do I handle missing values and outliers?
For missing values you can delete the observation, or replace the value with an estimate such as the mean or median. For outliers, first check whether they are errors, then remove, trim or winsorize them. Keep genuine extremes if they carry information.
What comes first, cleansing or scaling?
Cleansing comes first. You fix missing, invalid, inaccurate, non-uniform and duplicate data, and then you transform and scale the variables.