Identifying Exploratory Data Analysis in ML Pipelines
A company is building an ML model. The company collected new data and analyzed the data by creating a correlation matrix, calculating statistics, and visualizing the data. Which stage of the ML pipeline is the company currently in?
Community Votes
100% of anonymous learners picked answer C. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The exam tests the ability to distinguish between data exploration (understanding patterns) and data preparation (cleaning/transforming), with the trap being confusion between EDA and pre-processing steps.
The question tests recognition of Exploratory Data Analysis (EDA) through specific activities like correlation matrices and data visualization. Community consensus confirms that understanding data structure via statistics and visuals is the core definition of EDA.
Candidates often select 'Data pre-processing' because they associate statistical calculations with cleaning, but pre-processing focuses on handling missing values or scaling, whereas the prompt emphasizes analysis and visualization.
Community Discussion (3 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Exploratory Data Analysis (EDA) is defined by the process of investigating datasets to summarize their main characteristics, often using visual methods. Creating a correlation matrix helps identify relationships between variables, while calculating statistics and visualizing data reveals distributions and anomalies. These actions are purely analytical and aim to understand the data before any transformation occurs.Why the Other Options Are Wrong
Data pre-processing involves transforming raw data into a usable format, such as imputing missing values or normalizing scales, which is not described here. Feature engineering focuses on creating new attributes from existing ones to improve model performance, rather than just analyzing existing patterns. Hyperparameter tuning is a later stage where algorithm parameters are optimized to minimize error, unrelated to initial data inspection.Community Comment Notes
Comments unanimously support option C, highlighting that EDA is the phase for examining underlying structures and detecting patterns. One comment notes that these tasks inform later stages like preprocessing, reinforcing the chronological order of the ML pipeline.Official Reference
Exam Strategy
Look for keywords related to 'understanding,' 'visualizing,' and 'summarizing' to identify EDA. Distinguish this from 'cleaning' or 'transforming' which indicate pre-processing, and 'building' or 'optimizing' which indicate feature engineering or tuning.
Related Analysis
Practice All AIF-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full AIF-C01 Practice Test →