Skip to content

CFA Level I Exam · Introduction to Financial Data Science

Fintech and Data Science in Finance for CFA Level I

Updated 7 October 2026 · Fact-checked

Fintech is the use of new technology to deliver and improve financial services. Data science extracts insight from data using statistics, computing and machine learning. Big data is described by volume, velocity, variety and veracity. On the exam, match each use case to the right tool and spot the data limitation in the stem.

Understand Fintech and Data Science in Finance

Fintech means technology applied to financial services. In investment management it covers areas such as automated advice, algorithmic trading, faster analysis of large data sets, and distributed ledger technology. Data science is the field that pulls useful information out of data. It combines statistics, computer science and domain knowledge.

Traditional finance data is structured and small: prices, financial statements, economic series. Today analysts also use alternative data, such as satellite images, web traffic, social media posts, card transactions and sensor readings. Much of it is unstructured: text, images, audio. It does not fit neatly into rows and columns.

Big data is the label for data sets that are too large, fast or varied for traditional tools. The common description uses the four Vs: volume (the amount of data), velocity (the speed at which it arrives and must be processed), variety (the range of formats: structured, semi-structured, unstructured) and veracity (the reliability and quality of the data). Some sources add a fifth V, value. Data quality matters: more data is not always better data.

Data science is used to forecast returns, detect fraud, assess credit risk, read news and filings with text analytics, and automate trading and portfolio tasks. Machine learning finds patterns without being told the exact rule. Natural language processing turns text into data you can analyse.

There are limits. Models can overfit, data can be biased or noisy, results can be hard to explain, and a pattern found in the past may not repeat. Good use needs human judgement, data checks and clear governance. The 2027 curriculum adds material on financial data science, AI and large language models, so expect questions about both the uses and the cautions.

Key formulas to remember

Big data characteristics (the Vs)
Volume = how much; Velocity = how fast; Variety = what formats; Veracity = how reliable
Learn each V with a one-line test. Volume is size, velocity is speed of generation and processing, variety is mix of structured and unstructured forms, veracity is quality and trustworthiness.
Structured vs unstructured data
Structured = fixed fields (tables, prices); Unstructured = no fixed format (text, images, audio)
Alternative data is mostly unstructured and needs processing before analysis.
Use case matching
Text/news/filings → NLP; pattern finding and prediction → machine learning; fast rule-based order execution → algorithmic trading
Most application questions reduce to picking the right tool for the described task.

How to solve Fintech and Data Science in Finance questions

Use this method for any question on fintech, big data or data science applications.

  1. 1Read the stem and name the task: forecasting, classification, reading text, trading, advice, record keeping or fraud detection.
  2. 2Identify the data described: structured or unstructured, large or small, fast or slow, reliable or noisy.
  3. 3If the stem asks about big data, map the clue to a V: size means volume, speed means velocity, mixed formats mean variety, quality doubts mean veracity.
  4. 4Match the task to the tool: text goes to NLP, pattern detection to machine learning, rule-based execution to algorithms.
  5. 5Check for a limitation the stem hints at, such as overfitting, bias, poor data quality or lack of explainability.
  6. 6Eliminate the two options that name the wrong V or tool, or that overstate what the technology can do.
  7. 7Choose the option that is most specific to the facts in the stem.

Quickest way: Clue-to-V shortcut

When to use it: Use it when a question describes a data set and asks which big data characteristic it shows or which challenge it creates.

  1. Underline the key adjective in the stem: huge, real-time, mixed or unreliable.
  2. Convert it: huge = volume, real-time = velocity, mixed = variety, unreliable = veracity.
  3. Cross out options that name a different V.
  4. If two options remain, pick the one that fits the exact wording of the stem, not the general idea.

Common mistakes in Fintech and Data Science in Finance

  • Confusing velocity with volume

    Both sound like 'a lot of data', so students treat them as the same.

    Fix: Volume is how much data exists. Velocity is how quickly it is created and must be processed. Look for words like real-time or streaming.

  • Thinking big data means only structured data

    Students link finance data with prices and spreadsheets.

    Fix: Variety includes unstructured sources such as text, images and audio. Alternative data is mostly unstructured.

  • Assuming more data always means better forecasts

    Technology hype makes big data sound like a guarantee.

    Fix: Quality matters. Noisy, biased or unreliable data (a veracity problem) can produce poor or misleading models.

  • Treating fintech as one thing

    Students memorise the word without the sub-areas.

    Fix: Separate the uses: automated advice, algorithmic trading, data analysis, and distributed ledger technology. Then match each to its function.

  • Believing machine learning removes the need for human judgement

    Automation sounds complete.

    Fix: Models can overfit, inherit bias and be hard to explain. Analysts still validate inputs, outputs and fit with the investment process.

  • Mixing up a data type with a data source

    Terms like alternative data and unstructured data overlap.

    Fix: Structured and unstructured describe format. Alternative data describes where the data comes from, outside traditional financial sources.

Worked examples

Example 1

An asset manager collects tick-by-tick trade prices from many exchanges and must process the trade data within milliseconds. Which big data characteristic does the milliseconds requirement MOST clearly illustrate? A. Velocity. B. Veracity. C. Variety.

Show the solution
  1. The milliseconds processing requirement is about the speed at which data arrives and must be processed.
  2. Speed of generation and processing is the definition of velocity, so A fits.
  3. Veracity is about reliability and quality. The stem gives no concern about data quality, so B is wrong.
  4. Variety is about mixed formats. The milliseconds requirement says nothing about format, so C is wrong.

Answer: A. Velocity.

Example 2

An analyst wants to convert thousands of quarterly earnings call transcripts into a measure of management tone to use in a model. Which technology is MOST appropriate? A. Distributed ledger technology. B. Natural language processing. C. Rule-based order routing.

Show the solution
  1. The task is to turn text into numeric information.
  2. Natural language processing is built to analyse human language, so it fits the task.
  3. Distributed ledger technology is for shared record keeping, not text analysis.
  4. Rule-based order routing handles trade execution and does not read transcripts.

Answer: B. Natural language processing.

Exam tips

  • Expect short scenario questions: the stem gives a clue, and you name the V or the tool. Practise the clue-to-V conversion until it is automatic.
  • Watch for options that overstate technology, such as 'eliminates bias' or 'guarantees accuracy'. These are usually wrong.
  • Know the difference between structured and unstructured data and give a quick example of each.
  • With no penalty for wrong answers, never leave a question blank. Eliminate one option you can rule out and choose between the other two.
  • Read for limitations in the stem. A hint of poor data quality, overfitting or opaque models usually shapes the right answer.

Practice questions from Introduction to Financial Data Science

Fintech and Data Science in Finance in other exams

The same ground in other exams, if you are preparing for more than one or want another angle on it.

Fintech and Data Science in Finance: frequently asked questions

What are the characteristics of big data in CFA Level I?

The standard description is volume, velocity and variety, often with veracity added. Volume is the amount of data, velocity is the speed of generation and processing, variety is the mix of formats, and veracity is data reliability. Know a one-line meaning for each.

What is alternative data?

Alternative data comes from sources outside traditional company filings and market prices. Examples are satellite images, web traffic, social media, card transactions and sensor data. It is often unstructured and needs processing and quality checks.

How is fintech used in investment management?

Common uses include automated advice, algorithmic trading, faster analysis of large data sets, text analytics of news and filings, fraud detection and distributed ledger record keeping. Match each use to what it does when answering questions.

Do I need to calculate anything for this topic?

No. The topic is conceptual, so you will not need your TI BA II Plus or HP 12C. Questions test definitions, matching tools to tasks, and recognising limitations.