FRM Exam Part I · Machine-Learning Methods
Data Preparation and Feature Engineering for FRM Part I
Updated 11 October 2026 · Fact-checked
Data preparation turns raw data into inputs a model can learn from. You clean errors, treat missing values and outliers, scale features to comparable ranges, and select or create features. Fit every transformation on training data only, then apply it unchanged to validation and test data.
Understand Data Preparation and Feature Engineering
A machine-learning model is only as good as the data it sees. Raw risk data has typos, gaps, extreme values and columns on very different scales. Data preparation fixes these problems before training.
Cleaning means removing duplicates, correcting wrong entries and making formats consistent. Missing values can be handled by deleting rows, deleting columns, or imputing a value such as the mean, median or a model-based estimate. Deleting is simple but loses information and can bias results if data is not missing at random. The median is more robust than the mean when the data is skewed.
Outliers are extreme observations. They may be errors or real events, such as a crash day. Do not remove them blindly. Options are to investigate, cap them (winsorize), or transform the data, for example with logs. In risk work, genuine tail events are often the very thing you care about.
Scaling puts features on comparable ranges. This matters for models based on distances or penalties, such as k-nearest neighbors, clustering, PCA, and regularized regression like ridge and LASSO. Tree-based models are mostly insensitive to scaling. Standardization gives mean 0 and standard deviation 1. Min-max normalization maps values to the range 0 to 1.
Feature selection chooses a subset of existing features and drops the rest. Feature engineering creates new features from existing ones, such as a ratio, a lagged return or an interaction term. The biggest trap is data leakage: using information from the test set, or from the future, when preparing training data. Compute means, standard deviations, minimums and maximums on the training set only.
Key formulas to remember
- Standardization (z-score)
- z = (x − μ) ÷ σ
- μ and σ come from the training set. Result has mean 0 and standard deviation 1. It does not make the data normal.
- Min-max normalization
- x' = (x − min) ÷ (max − min)
- Maps training data to 0 to 1. Very sensitive to outliers because min and max are extreme values.
- Mean imputation
- x_missing = (Σ x_observed) ÷ n_observed
- Keeps the mean unchanged but shrinks variance. Use the median for skewed data.
- Interquartile range outlier rule
- Outlier if x < Q1 − 1.5 × IQR or x > Q3 + 1.5 × IQR, where IQR = Q3 − Q1
- A common rule of thumb, not a law. Flag for review rather than delete automatically.
How to solve Data Preparation and Feature Engineering questions
Use this sequence for any question on preparing data for a model.
- 1Identify the problem in the data: missing values, outliers, mixed scales, too many features, or leakage.
- 2Identify the model type. Distance-based and penalized models need scaling; trees usually do not.
- 3Choose the matching fix: impute or delete for missing data, investigate or cap for outliers, standardize or normalize for scale, select or engineer for features.
- 4Check which statistics are used. Scaling and imputation values must come from the training set only.
- 5If a calculation is needed, apply the formula with the training mean, standard deviation, min or max.
- 6Test each answer option against the risk of leakage, information loss or distortion, and pick the one that avoids them.
Quickest way: Scale, split, leakage check
When to use it: Use for conceptual multiple-choice questions where you must pick the correct preprocessing practice.
- Ask: does the model use distances or penalties? If yes, scaling is needed.
- Ask: were statistics computed before the train-test split? If yes, it is leakage and the option is wrong.
- Ask: is there an outlier? Prefer median or robust methods over mean, and prefer min-max only if outliers are absent.
- Selecting means dropping existing columns. Engineering means creating new ones.
Common mistakes in Data Preparation and Feature Engineering
Scaling the whole dataset before splitting into training and test sets
It is quicker and looks harmless.
Fix: Compute μ and σ (or min and max) on training data only, then apply them to the test data.
Believing standardization makes data normally distributed
The z-score formula is the same as the one for the normal distribution.
Fix: Standardization only shifts and rescales. The shape of the distribution, including skewness, is unchanged.
Deleting every outlier automatically
Outliers look like errors that distort averages.
Fix: Investigate first. Real tail events matter in risk models. Cap, transform or keep them if they are genuine.
Using min-max scaling on data with extreme outliers
Students forget min and max are themselves extremes.
Fix: One huge value compresses all other values into a tiny range. Use standardization or treat outliers first.
Mixing up feature selection and feature engineering
Both sound like improving the inputs.
Fix: Selection keeps a subset of existing features. Engineering builds new features from them.
Assuming mean imputation is always safe
It keeps the sample size and the mean.
Fix: It reduces variance and weakens correlations. It can also bias results if data is not missing at random.
Worked examples
Example 1
A training set of a feature has mean 8% and standard deviation 4%. A new observation in the test set is 14%. Standardize it for the model.
Show the solution
- Formula: z = (x − μ) ÷ σ.
- Use training statistics: μ = 8%, σ = 4%.
- z = (14 − 8) ÷ 4 = 6 ÷ 4 = 1.5.
Answer: z = 1.5. The observation is 1.5 training standard deviations above the training mean.
Example 2
A training feature has minimum 20, maximum 120 and one value of 70. The test set contains a value of 140. Apply min-max normalization to both values. What does the result show?
Show the solution
- Formula: x' = (x − min) ÷ (max − min), with training min = 20 and max = 120.
- Range = 120 − 20 = 100.
- For 70: x' = (70 − 20) ÷ 100 = 0.5.
- For 140: x' = (140 − 20) ÷ 100 = 1.2.
- The test value falls outside 0 to 1 because it exceeds the training maximum.
Answer: 70 becomes 0.5 and 140 becomes 1.2. Values outside the training range can give results beyond 0 to 1, and you must not recompute min and max using test data.
Exam tips
- Whenever a question mentions test or validation data, check for data leakage first. It is a favorite trap.
- Know which models need scaling: KNN, clustering, PCA, ridge and LASSO. Trees and random forests generally do not.
- Standardization calculations are short. Do them carefully using training mean and standard deviation, and watch the sign.
- Expect wording that contrasts selection (drop features) with engineering (create features). Match the definition exactly.
Practice questions from Machine-Learning Methods
- A bank compares a neural network with a regularized linear model for predicting corporate defaults using a modest dataset of 800 firms and 1…
- A risk analyst fits a regression of credit spread changes on 40 candidate predictors using only 120 observations. The ordinary least squares…
- A model predicting loan losses achieves a mean squared error of 0.5 on the training sample but 4.0 on a held-out validation sample. Which in…
- A network has 4 inputs, one hidden layer with 5 neurons, and 1 output neuron. Every neuron in a layer is connected to every neuron in the pr…
- A LASSO model minimizes the sum of squared residuals plus lambda times the sum of absolute coefficients. A model has coefficients 2.0, -1.5,…
Data Preparation and Feature Engineering in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Data Preparation and Feature Engineering: frequently asked questions
What is the difference between standardization and normalization?
Standardization subtracts the mean and divides by the standard deviation, giving mean 0 and standard deviation 1. Normalization (min-max) rescales values to a fixed range, usually 0 to 1. Normalization is more affected by outliers.
Why must scaling be fitted on training data only?
Using test data to compute the mean or standard deviation lets information from the test set leak into training. That makes performance look better than it will be on new data.
What is the difference between feature selection and feature engineering?
Feature selection chooses a subset of existing features and removes the rest. Feature engineering creates new features, such as ratios, lags or interaction terms, from the existing data.
How should missing values be handled?
It depends on how much is missing and why. You can delete rows or columns, or impute with the mean, median or a model estimate. Prefer the median for skewed data and be cautious if data is not missing at random.