Skip to content

CFA Level I · CFA Level I Exam

Introduction to Financial Data Science for CFA Level I

Financial data science applies data tools, machine learning and text analysis to finance problems. For Level I, you must know data types, data preparation, supervised and unsupervised learning, overfitting, NLP basics and chart choice. Questions are mostly conceptual: match the method or problem to the right definition, then eliminate two options.

What this chapter covers

This chapter introduces the tools modern finance uses to turn raw data into decisions. It starts with fintech and the data science workflow, then covers data types and cleaning, machine learning approaches and model fit, and finally text analytics, natural language processing (NLP) and visualization. It is new in the 2027 curriculum, which CFA Institute says includes a new learning module on financial data science, AI and large language models.

The chapter is mostly conceptual. You will not do long calculations. You will be asked to name things correctly: structured versus unstructured data, supervised versus unsupervised learning, overfitting versus underfitting, tokenization versus stemming. Precise vocabulary is what earns marks.

It connects to the rest of the paper in several places. Quantitative Methods gives the statistics and regression ideas behind model fit. Portfolio Construction, Equities and Fixed Income are areas where data-driven tools are applied. Ethics matters too, because data use raises issues such as privacy, bias and diligence in using outputs you do not fully understand.

Because this content is new, many candidates will have less practice with it, and questions are usually short, definition-led and quick to answer if you know the terms. Each question is worth the same as any other, and there is no penalty for a wrong answer, so every item you can answer in under 90 seconds buys time for harder numerical questions elsewhere. Clear vocabulary here is a cheap, reliable gain.

Introduction to Financial Data Science: topics in the order to study them

  1. 1Fintech and Data Science in FinanceStart here to get the big picture and the common vocabulary of the field before details.
  2. 2Data Types and Data PreparationModels depend on data, so learn what data looks like and how it is cleaned before learning the models.
  3. 3Machine Learning Approaches and Model FitThis builds on data concepts and is the core of the chapter, so study it once the inputs are clear.
  4. 4Text Analytics, NLP and Data VisualizationLast, because text analysis applies machine learning to unstructured data, and visualization ties results together.

How to prepare Introduction to Financial Data Science

Treat this as a vocabulary and classification chapter. Your aim is to recognise a scenario and name the right concept fast.

  1. Read each topic once for the story: data comes in, is cleaned, is modelled, and results are communicated.
  2. Build a one-page glossary in your own words. Add a one-line example from finance for every term.
  3. Make comparison pairs: structured vs unstructured, supervised vs unsupervised, overfitting vs underfitting, training vs test data.
  4. For each scenario you read, ask: what is the data type, what is the goal, and is there a labelled target? This usually points to the method.
  5. Do short question sets right after each topic, then a mixed set. For each miss, write why the other two options were wrong.
  6. Revise weekly with flashcards on your phone. Ten minutes a day works well for definition-heavy material.
  7. Before the exam, rehearse elimination: remove the option that mixes up two terms, then choose between the remaining two.

Common mistakes in Introduction to Financial Data Science

  • Mixing up supervised and unsupervised learning

    Fix: Ask one question: is there a known target the model is trained to predict? If yes, it is supervised.

  • Confusing overfitting with underfitting

    Fix: Link overfitting to great training results but weak out-of-sample results. Underfitting is weak everywhere.

  • Treating text as already structured

    Fix: Remember that text is unstructured until NLP steps such as tokenization convert it into a usable form.

  • Choosing a regression or classification label by the input instead of the output

    Fix: Look at the target: a number means regression, a category means classification.

  • Skipping this chapter because it looks non-technical or new

    Fix: Give it steady, short study sessions. Questions are quick to answer if you know the terms.

  • Learning definitions without finance examples

    Fix: Attach one example to each term, such as credit default prediction for classification or sentiment scoring for NLP.

Last-day revision: Introduction to Financial Data Science

  • Structured data fits fixed fields and tables; unstructured data such as text, images and audio does not.
  • Supervised learning uses labelled data to predict a target; unsupervised learning finds patterns without labels.
  • Regression predicts a continuous target; classification predicts a category.
  • Clustering and dimension reduction are typical unsupervised tasks.
  • Overfitting means the model fits training data too closely and performs poorly on new data.
  • Underfitting means the model is too simple to capture the pattern, even in training data.
  • Split data into training, validation and test sets; the test set checks performance on unseen data.
  • Data cleaning deals with missing values, outliers, duplicates and inconsistent formats.
  • Tokenization splits text into units; stop-word removal and stemming or lemmatization reduce noise.
  • NLP turns text into structured form so it can be analysed, for example for sentiment.
  • Pick the chart for the message: trends over time, comparisons, distributions or relationships.
  • Always apply Ethics thinking: data privacy, bias and understanding a model before relying on it.

Introduction to Financial Data Science practice questions

Introduction to Financial Data Science in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Introduction to Financial Data Science: frequently asked questions

Does this chapter need maths or calculator work?

Very little. It is mainly conceptual, so you need clear definitions and the ability to match scenarios to methods. Save your calculator practice for Quantitative Methods and other numerical chapters.

Do I need programming knowledge for Level I?

No. You are tested on concepts, uses and limitations of the tools, not on writing code. Know what each technique does and when it fits.

What order should I study the topics in?

Follow the chapter order: fintech and data science overview, data types and preparation, machine learning and model fit, then text analytics, NLP and visualization. Each step builds on the one before.

How are questions in this chapter asked?

Like every Level I item, they are standalone with three options (A, B, C). Expect short scenarios asking you to identify a data type, a learning approach or a model-fit problem.