Which AWS solution detects changing schemas and loads data to S3 with least overhead?
A company extracts approximately 1 TB of data every day from data sources such as SAP HANA, Microsoft SQL Server, MongoDB, Apache Kafka, and Amazon DynamoDB. Some of the data sources have undefined data schemas or data schemas that change. A data engineer must implement a solution that can detect the schema for these data sources. The solution must extract, transform, and load the data to an Amazon S3 bucket. The company has a service level agreement (SLA) to load the data into the S3 bucket within 15 minutes of data creation. Which solution will meet these requirements with the LEAST operational overhead?
Community Votes
100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
It tests selecting a managed ETL service that performs schema inference across heterogeneous sources; the trap is proposing self-managed Apache Spark on EMR or Lambda for a 1 TB/day pipeline with a 15-minute SLA.
This DEA-C01 scenario requires detecting undefined or changing schemas across SAP HANA, SQL Server, MongoDB, Kafka, and DynamoDB while loading transformed data into Amazon S3 within 15 minutes. AWS Glue is the managed choice that satisfies the schema detection and low-operational-overhead requirements.
Choosing Amazon EMR with Apache Spark because it can detect schemas and process big data, but self-managed clusters add the operational overhead that AWS Glue eliminates.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
AWS Glue is a serverless, managed ETL service, so it directly addresses the LEAST operational overhead requirement while still detecting schemas for undefined or changing sources. Glue crawlers can infer schemas from SAP HANA, Microsoft SQL Server, MongoDB, Kafka, and DynamoDB through connectors, and Glue jobs run Apache Spark without clusters to manage. For approximately 1 TB per day and a 15-minute SLA, scheduled Glue jobs and job bookmarks can process incremental data reliably. Because Glue handles both schema inference and the extract, transform, and load steps into Amazon S3, option B matches every stated requirement.Why the Other Options Are Wrong
Option A uses Amazon EMR with Apache Spark, which can detect schemas and process large volumes, but it requires provisioning, tuning, and maintaining clusters, adding unnecessary operational overhead. Option C puts PySpark in AWS Lambda, which has a 15-minute maximum execution time and packaging limits; a 1 TB daily ETL pipeline with schema detection is impractical in Lambda and offers no managed schema discovery. Option D uses an Amazon Redshift stored procedure and Redshift Spectrum, but Redshift Spectrum queries data in S3 rather than extracting from sources like Kafka or DynamoDB, and a stored procedure is not a schema-detection or ingestion mechanism for those systems.Community Comment Notes
The community consensus is strongly aligned with B: rralucard_ selected the AWS Glue option, and GiorgioGss wrote, "Least effort = B" to capture the operational-overhead reasoning. kj07 also stated that "the least operational overhead is B," while Christina666 summarized the decision as "Glue ETL." These comments reinforce that the exam is testing recognition of a managed Glue pipeline, not a self-managed Spark or Lambda implementation.Official Reference
Exam Strategy
When a DEA-C01 question emphasizes LEAST operational overhead and schema detection across multiple sources, favor managed AWS Glue over self-managed EMR or custom code. Watch for 15-minute SLAs and 1 TB/day volumes to eliminate Lambda and Redshift stored procedures.
Frequently Asked Questions
Why is AWS Glue better than Amazon EMR for schema detection here?
Glue is serverless, so it infers changing schemas with crawlers and runs Apache Spark ETL without clusters to provision or tune, which directly meets the least-operational-overhead requirement.
Can AWS Lambda handle 1 TB of daily ETL within a 15-minute SLA?
No. Lambda has a 15-minute maximum runtime and limited packaging, making it impractical for a 1 TB daily pipeline with schema detection across SAP HANA, Kafka, and DynamoDB.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →