Which AWS solution detects changing schemas and loads data to S3 with least overhead?

Perform data ingestion. Transform and process data. Design data models and schema evolution.
Answer Correct answer: B — Use AWS Glue to detect schemas and run Spark ETL into Amazon S3 for changing or undefined schemas at low operational overhead.

A company extracts approximately 1 TB of data every day from data sources such as SAP HANA, Microsoft SQL Server, MongoDB, Apache Kafka, and Amazon DynamoDB. Some of the data sources have undefined data schemas or data schemas that change. A data engineer must implement a solution that can detect the schema for these data sources. The solution must extract, transform, and load the data to an Amazon S3 bucket. The company has a service level agreement (SLA) to load the data into the S3 bucket within 15 minutes of data creation. Which solution will meet these requirements with the LEAST operational overhead?

  1. Use Amazon EMR to detect the schema and to extract, transform, and load the data into the S3 bucket. Create a pipeline in Apache Spark.
  2. Use AWS Glue to detect the schema and to extract, transform, and load the data into the S3 bucket. Create a pipeline in Apache Spark. Correct Answer
  3. Create a PySpark program in AWS Lambda to extract, transform, and load the data into the S3 bucket.
  4. Create a stored procedure in Amazon Redshift to detect the schema and to extract, transform, and load the data into a Redshift Spectrum table. Access the table from Amazon S3.

Community Votes

B
100%

100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

It tests selecting a managed ETL service that performs schema inference across heterogeneous sources; the trap is proposing self-managed Apache Spark on EMR or Lambda for a 1 TB/day pipeline with a 15-minute SLA.

This DEA-C01 scenario requires detecting undefined or changing schemas across SAP HANA, SQL Server, MongoDB, Kafka, and DynamoDB while loading transformed data into Amazon S3 within 15 minutes. AWS Glue is the managed choice that satisfies the schema detection and low-operational-overhead requirements.

Choosing Amazon EMR with Apache Spark because it can detect schemas and process big data, but self-managed clusters add the operational overhead that AWS Glue eliminates.

Community Discussion (4 comments)

GiorgioGss 👍 6 Selected: B
Least effort = B
rralucard_ 👍 5 Selected: B
B. Use AWS Glue to detect the schema and to extract, transform, and load the data into the S3 bucket. Create a pipeline in Apache Spark.
Christina666 👍 1 Selected: B
Glue ETL
kj07 👍 3
The option with the least operational overhead is B.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

AWS Glue is a serverless, managed ETL service, so it directly addresses the LEAST operational overhead requirement while still detecting schemas for undefined or changing sources. Glue crawlers can infer schemas from SAP HANA, Microsoft SQL Server, MongoDB, Kafka, and DynamoDB through connectors, and Glue jobs run Apache Spark without clusters to manage. For approximately 1 TB per day and a 15-minute SLA, scheduled Glue jobs and job bookmarks can process incremental data reliably. Because Glue handles both schema inference and the extract, transform, and load steps into Amazon S3, option B matches every stated requirement.

Why the Other Options Are Wrong

Option A uses Amazon EMR with Apache Spark, which can detect schemas and process large volumes, but it requires provisioning, tuning, and maintaining clusters, adding unnecessary operational overhead. Option C puts PySpark in AWS Lambda, which has a 15-minute maximum execution time and packaging limits; a 1 TB daily ETL pipeline with schema detection is impractical in Lambda and offers no managed schema discovery. Option D uses an Amazon Redshift stored procedure and Redshift Spectrum, but Redshift Spectrum queries data in S3 rather than extracting from sources like Kafka or DynamoDB, and a stored procedure is not a schema-detection or ingestion mechanism for those systems.

Community Comment Notes

The community consensus is strongly aligned with B: rralucard_ selected the AWS Glue option, and GiorgioGss wrote, "Least effort = B" to capture the operational-overhead reasoning. kj07 also stated that "the least operational overhead is B," while Christina666 summarized the decision as "Glue ETL." These comments reinforce that the exam is testing recognition of a managed Glue pipeline, not a self-managed Spark or Lambda implementation.

Official Reference

Exam Strategy

When a DEA-C01 question emphasizes LEAST operational overhead and schema detection across multiple sources, favor managed AWS Glue over self-managed EMR or custom code. Watch for 15-minute SLAs and 1 TB/day volumes to eliminate Lambda and Redshift stored procedures.

Frequently Asked Questions

Why is AWS Glue better than Amazon EMR for schema detection here?

Glue is serverless, so it infers changing schemas with crawlers and runs Apache Spark ETL without clusters to provision or tune, which directly meets the least-operational-overhead requirement.

Can AWS Lambda handle 1 TB of daily ETL within a 15-minute SLA?

No. Lambda has a 15-minute maximum runtime and limited packaging, making it impractical for a 1 TB daily pipeline with schema detection across SAP HANA, Kafka, and DynamoDB.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide