Determining Whether a Dataset Can Predict Its Target Variable Using SageMaker Autopilot
A company uses Amazon Athena to query a dataset in Amazon S3. The dataset has a target variable that the company wants to predict. The company needs to use the dataset in a solution to determine if a model can predict the target variable. Which solution will provide this information with the LEAST development effort?
Community Votes
100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
SageMaker Autopilot automates the entire model development process, including preprocessing, model selection, hyperparameter tuning, and performance evaluation, and reports the achieved performance, so it answers the predictability question without hand-written modeling code.
A company queries a dataset in S3 with Athena and wants to know whether a model can predict the target variable in that dataset, with the least development effort. What is needed is not a specific model but a fast, automated answer about whether the target is predictable from the available features.
Configuring Amazon Macie to analyze the dataset, which is a sensitive data discovery and classification service with no model building or evaluation capability. Also wrong is hand-rolling linear regression on EC2, which is explicitly the high-effort path the question rules out.
Community Discussion (3 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The question asks whether the target variable can be predicted from the dataset, with the least development effort, and SageMaker Autopilot directly answers that. Autopilot automates the model development process end to end, taking the dataset, exploring it, preprocessing the features, training candidate models, tuning them, and evaluating their performance, and it reports the achieved performance metrics. Because the target is numeric and the dataset is tabular, Autopilot is the appropriate tool, and it produces the performance evidence needed to judge predictability without writing any modeling code. The vote was unanimous at 100 for A, and chris_spencer called A the most logical answer while explaining that the alternatives all involve far more setup effort.Why the Other Options Are Wrong
Implementing custom scripts for preprocessing, linear regression, and evaluation on EC2 instances (B) is precisely the high development effort the question excludes, since the engineer would have to build and maintain every step of the modeling pipeline manually. Configuring Amazon Macie to analyze the dataset and create a model (C) is a category error, because Macie is a managed data discovery and classification service for finding sensitive information and has no model training or performance evaluation capability at all. Selecting a model from Amazon Bedrock and tuning it (D) also carries substantial effort and, as Sadrik pointed out, Bedrock supplies foundation models for language tasks such as text generation rather than structured tabular prediction, so it is the wrong tool for a numeric target.Community Comment Notes
The community was unanimous at 100 for A. chris_spencer evaluated all four options and found A the most logical, citing the excessive setup effort of the custom scripts on EC2, the non-fit of Macie, and the effort of standing up Bedrock. Sadrik supplied the strongest technical objection to option D, explaining that Bedrock provides foundation models for text generation and is not used for structured data prediction, which is the same reasoning that makes Autopilot the right choice for a tabular target.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →