CFA Level II Exam · Big Data Projects
Big Data Characteristics and Project Workflow for CFA Level II
Updated 7 October 2026 · Fact-checked
Big data is defined by volume, velocity, variety and veracity. A big data modeling project follows set steps: conceptualize the task, collect data, prepare and wrangle it, explore it, then train and evaluate the model. Exam questions ask you to match a described situation to a characteristic or a workflow step.
Understand Big Data Characteristics and Project Workflow
Big data means datasets so large, fast or varied that traditional tools struggle with them. Investment firms use it to find signals in sources like social media posts, satellite images, web traffic and transaction records, alongside conventional financial data.
The curriculum describes big data with the four Vs. Volume is the quantity of data. Velocity is how fast data arrives and must be processed, from batch to real time. Variety is the range of formats and sources. Veracity is the reliability and credibility of the data. Veracity matters because a huge dataset full of noise, bias or fake content can mislead a model.
Variety links to the data types. Structured data fits a fixed schema of rows and columns, like a table of prices or financial ratios. Unstructured data has no such organization, like text, images, audio and video. Semi-structured data sits between the two: it has some organizing tags or fields but is not a strict table, for example HTML pages or JSON files. Unstructured data usually needs extra processing, such as converting text into numbers, before a model can use it.
A big data project follows a workflow. For structured data the steps are: conceptualization of the modeling task, data collection, data preparation and wrangling, data exploration, and model training. For unstructured text the curriculum uses a parallel set: text problem formulation, data curation, text preparation and wrangling, text exploration, and model training. In both cases, preparation (cleansing and preprocessing) and exploration come before training, and model training includes evaluating performance and tuning.
The logic is simple. You must define the problem first, get the right data, clean it, understand it, and only then fit and test a model. Skipping early steps produces a model that is accurate on paper but useless in practice.
Key formulas to remember
- Four Vs of big data
- Volume (quantity) | Velocity (speed of arrival) | Variety (formats and sources) | Veracity (reliability)
- Match each exam description to one V. Credibility, noise and bias point to veracity.
- Structured-data workflow
- Conceptualization → Data collection → Data preparation and wrangling → Data exploration → Model training
- Order matters. Exploration comes before training, and training includes evaluation and tuning.
- Text-data workflow
- Text problem formulation → Data curation → Text preparation and wrangling → Text exploration → Model training
- This is the unstructured text counterpart of the structured workflow.
- Data types
- Structured (fixed schema) | Semi-structured (tags or fields, no strict table) | Unstructured (no organization)
- Text, images and video are unstructured. Convert them to structured form before modeling.
How to solve Big Data Characteristics and Project Workflow questions
Use this method for any item-set question on big data characteristics or the project workflow.
- 1Read the question first, then find the part of the vignette that describes the data or the activity.
- 2Decide what is being asked: a characteristic (a V), a data type, or a workflow step.
- 3For a characteristic, find the key word: size means volume, speed means velocity, mixed formats means variety, trustworthiness means veracity.
- 4For a data type, ask whether the data fits rows and columns with a fixed schema. If not, it is unstructured or semi-structured.
- 5For a workflow step, name the activity: defining the goal is conceptualization, gathering is collection, cleaning is preparation, inspecting is exploration, fitting and testing is training.
- 6Check the order of steps. A step that should happen earlier cannot be the right answer for a later stage.
- 7Eliminate options that mix up neighboring concepts, such as veracity with variety, or preparation with exploration.
Quickest way: Keyword matching
When to use it: Use when you have under a minute per question and the vignette clearly describes the data or the activity.
- Underline one key word in the vignette: size, speed, formats, reliability, cleaning, inspecting, fitting.
- Map it: size = volume, speed = velocity, formats = variety, reliability = veracity.
- For workflow, map it: goal = conceptualization, gathering = collection, cleaning = preparation, inspecting = exploration, fitting = training.
- Pick the option that matches, and skip the rest.
Common mistakes in Big Data Characteristics and Project Workflow
Confusing variety with veracity.
Both words sound like quality measures and both appear in the same list.
Fix: Variety is about different formats and sources. Veracity is about whether the data can be trusted.
Treating velocity as the same as volume.
Fast data tends to produce a lot of data, so the two blur together.
Fix: Volume is how much data exists. Velocity is how quickly it arrives and must be processed.
Placing data exploration after model training.
Students think of exploring results rather than exploring inputs.
Fix: Exploration examines the prepared data to understand patterns and select features. It comes before training.
Mixing up data preparation and data exploration.
Both work with the dataset and both involve looking at it.
Fix: Preparation cleans and transforms data (fixing errors, handling missing values). Exploration investigates it (summaries, visuals, feature selection).
Calling all text data structured because it can be stored in a file.
Storage in a database looks like organization.
Fix: Structured means a fixed schema of fields. Free text, images and video are unstructured until converted.
Forgetting that model evaluation belongs to the training step.
Students expect evaluation to be its own stage.
Fix: In the curriculum workflow, model training includes performance evaluation and tuning.
Worked examples
Example 1
A fund collects millions of social media posts every hour. The posts include text, photos and short videos. Some come from automated accounts posting false claims about listed companies. Q1: Which characteristic best describes the arrival of posts every hour? Q2: Which characteristic does the bot content threaten? Q3: Is the data mostly structured or unstructured?
Show the solution
- Q1: The key idea is how fast data arrives and must be handled. That is velocity.
- Q2: False claims from bots reduce reliability and credibility. That is veracity.
- Q3: Text, photos and video have no fixed row-and-column schema, so they are unstructured.
Answer: Q1: Velocity. Q2: Veracity. Q3: Unstructured.
Example 2
An analyst plans to predict loan defaults. She first writes down the prediction goal and the target variable. She then gathers borrower records, removes duplicate rows and fills missing income values, plots distributions to pick useful features, and finally fits a model and checks it on held-out data. Q1: What is the first step she completed? Q2: Which step did removing duplicates and filling missing values belong to? Q3: Which step includes checking the model on held-out data?
Show the solution
- Q1: Defining the goal and target variable is the conceptualization of the modeling task.
- Q2: Removing duplicates and handling missing values is cleaning and transforming the data, which is data preparation and wrangling.
- Q3: Plotting distributions to choose features was data exploration, which comes before training. Fitting the model and checking it on held-out data is model training, which includes evaluation.
Answer: Q1: Conceptualization. Q2: Data preparation and wrangling. Q3: Model training (including evaluation).
Exam tips
- Read the question stem before the vignette so you know whether to look for a V, a data type or a workflow step.
- Reduce each V to one word in your head: size, speed, formats, trust.
- Memorize the workflow order for both structured and text data. Many questions are just ordering checks.
- When two workflow options seem right, pick the one whose activity is cleaning (preparation) versus inspecting (exploration).
- Do not leave any question blank. There is no penalty for wrong answers.
Big Data Characteristics and Project Workflow in other exams
The same ground in other exams, if you are preparing for more than one or want another angle on it.
Big Data Characteristics and Project Workflow: frequently asked questions
What are the four Vs of big data in CFA Level II?
They are volume, velocity, variety and veracity. Volume is size, velocity is speed of data arrival, variety is the range of formats and sources, and veracity is reliability.
What is the difference between structured and unstructured data?
Structured data fits a fixed schema of rows and columns, such as a table of prices. Unstructured data, such as text, images and video, has no such format. Semi-structured data has some tags or fields but no strict table.
What are the steps in a big data project workflow?
For structured data: conceptualization, data collection, data preparation and wrangling, data exploration, and model training. For text data the equivalent steps are problem formulation, data curation, text preparation and wrangling, text exploration, and model training.
Does model evaluation count as a separate step?
In the curriculum workflow, evaluation is part of model training. Training covers fitting the model, assessing its performance and tuning it.