Identifying Exploratory Data Analysis in ML Pipelines

Machine Learning Fundamentals

A company is building an ML model. The company collected new data and analyzed the data by creating a correlation matrix, calculating statistics, and visualizing the data. Which stage of the ML pipeline is the company currently in?

  1. Data pre-processing
  2. Feature engineering
  3. Exploratory data analysis Source Reference Answer
  4. Hyperparameter tuning

Community Votes

C
100%

100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The exam tests the ability to distinguish between data exploration (understanding patterns) and data preparation (cleaning/transforming), with the trap being confusion between EDA and pre-processing steps.

The question tests recognition of Exploratory Data Analysis (EDA) through specific activities like correlation matrices and data visualization. Community consensus confirms that understanding data structure via statistics and visuals is the core definition of EDA.

Candidates often select 'Data pre-processing' because they associate statistical calculations with cleaning, but pre-processing focuses on handling missing values or scaling, whereas the prompt emphasizes analysis and visualization.

Community Discussion (3 comments)

Jessiii 👍 1 Selected: C
C. Exploratory data analysis (EDA): EDA is the process of analyzing and visualizing the data to understand its characteristics and identify patterns, relationships, and anomalies. The company is performing actions like creating a correlation matrix, calculating statistics, and visualizing the data, all of which are typical activities in EDA.
Moon 👍 2 Selected: C
C: Exploratory data analysis Explanation: Exploratory Data Analysis (EDA) involves examining and summarizing data to understand its underlying structure, detect patterns, identify relationships (e.g., via a correlation matrix), and highlight any anomalies. The company's activities, such as creating a correlation matrix, calculating statistics, and visualizing the data, are typical tasks performed during EDA. Why not the other options? A: Data pre-processing: Data pre-processing involves cleaning and preparing data for modeling, such as handling missing values, scaling features, or encoding categorical data. While pre-processing may follow EDA, the tasks described in the question focus on analysis rather than preparation.
dehkon 👍 2
C. Exploratory data analysis Exploratory Data Analysis (EDA) involves examining and visualizing data to understand its structure, patterns, and relationships. Creating a correlation matrix, calculating statistics, and visualizing data are all typical tasks during the EDA phase, which helps inform later stages such as data preprocessing and feature engineering.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Exploratory Data Analysis (EDA) is defined by the process of investigating datasets to summarize their main characteristics, often using visual methods. Creating a correlation matrix helps identify relationships between variables, while calculating statistics and visualizing data reveals distributions and anomalies. These actions are purely analytical and aim to understand the data before any transformation occurs.

Why the Other Options Are Wrong

Data pre-processing involves transforming raw data into a usable format, such as imputing missing values or normalizing scales, which is not described here. Feature engineering focuses on creating new attributes from existing ones to improve model performance, rather than just analyzing existing patterns. Hyperparameter tuning is a later stage where algorithm parameters are optimized to minimize error, unrelated to initial data inspection.

Community Comment Notes

Comments unanimously support option C, highlighting that EDA is the phase for examining underlying structures and detecting patterns. One comment notes that these tasks inform later stages like preprocessing, reinforcing the chronological order of the ML pipeline.

Official Reference

Exam Strategy

Look for keywords related to 'understanding,' 'visualizing,' and 'summarizing' to identify EDA. Distinguish this from 'cleaning' or 'transforming' which indicate pre-processing, and 'building' or 'optimizing' which indicate feature engineering or tuning.

Related Analysis

Practice All AIF-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full AIF-C01 Practice Test →

← Back to AIF-C01 Study Guide