Amazon EMR for Petabyte-Scale Migration with Legacy Frameworks

Answer Correct answer: B — Amazon EMR meets requirements by supporting all listed frameworks (Pig, Oozie, Spark, HBase, Flink) and offering EMR Serverless to reduce operational overhead for petabyte-scale data processing.

A company is migrating on-premises workloads to AWS. The company wants to reduce overall operational overhead. The company also wants to explore serverless options. The company's current workloads use Apache Pig, Apache Oozie, Apache Spark, Apache Hbase, and Apache Flink. The on-premises workloads process petabytes of data in seconds. The company must maintain similar or better performance after the migration to AWS. Which extract, transform, and load (ETL) service will meet these requirements?

  1. AWS Glue
  2. Amazon EMR Correct Answer
  3. AWS Lambda
  4. Amazon Redshift

Community Votes

B
82%
A
18%

82% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The exam tests the ability to distinguish between fully managed serverless ETL (Glue) and managed big data platforms (EMR) when legacy framework compatibility is a strict requirement.

This question evaluates the selection of an AWS service for migrating complex, petabyte-scale ETL workloads involving Apache Pig, Oozie, Spark, HBase, and Flink. Amazon EMR is the correct choice because it natively supports these specific open-source frameworks while offering a serverless deployment option to reduce operational overhead.

Many candidates choose AWS Glue (A) because it is explicitly described as a 'serverless ETL' service in the question stem, overlooking the fact that Glue does not support the specific legacy frameworks listed (Pig, Oozie, HBase).

Community Discussion (19 comments)

milofficial 👍 18 Selected: B
Glue is like the more good-looking one, but weaker brother of EMR. So when it's about petabyte scales, let EMR do the work and have Glue stay away from the action.
Ell89 👍 1 Selected: B
Glue doesnt natively support Pig, HBase and Flink.
Udyan 👍 1 Selected: B
Apache = EMR
heavenlypearl 👍 2 Selected: B
Amazon EMR Serverless is a deployment option for Amazon EMR that provides a serverless runtime environment. This simplifies the operation of analytics applications that use the latest open-source frameworks, such as Apache Spark and Apache Hive. With EMR Serverless, you don’t have to configure, optimize, secure, or operate clusters to run applications with these frameworks. https://docs.aws.amazon.com/emr/latest/EMR-Serverless-UserGuide/emr-serverless.html
87ebc7d 👍 2
Discarded, not 'discarted'. 'Discarted' isn't a word.
leotoras 👍 1
B. Amazon EMR Serverless is a deployment option for Amazon EMR that provides a serverless runtime environment. This simplifies the operation of analytics applications that use the latest open-source frameworks, such as Apache Spark and Apache Hive. With EMR Serverless, you don’t have to configure, optimize, secure, or operate clusters to run applications with these frameworks.
Eleftheriia 👍 2 Selected: A
I think it is A, Glue • Amazon EMR is used for petabyte-scale data collection and data processing. • AWS Glue is used as a serverless and managed ETL service, and also used for managing data quality with AWS Glue Data Quality.
San_Juan 👍 1 Selected: A
Glue. It talks about "serverless" so EMR is discarted. The mention of Spark, Hbase, etc is for confusing you, because it doesn't say that they wanted to keep using them. Glue can run Spark using "glueContext" (similar a SparkContext) for reading tables, files and create frames.
sachin 👍 1
The company also wants to explore serverless options. ? Glue (A). or EMR Serverless
V0811 👍 1 Selected: A
Serverless: AWS Glue is a fully managed, serverless ETL service that automates the process of data discovery, preparation, and transformation, helping minimize operational overhead.Integration with Big Data Tools: It integrates well with various AWS services and supports Spark jobs for ETL purposes, which aligns well with Apache Spark workloads.Performance: AWS Glue can handle large-scale ETL workloads, and it is designed to manage petabytes of data efficiently, comparable to the performance of on-premises solutions.While B. Amazon EMR could also be considered for its flexibility in handling big data workloads using tools like Apache Spark, it requires more management and doesn't fit the serverless requirement as closely as AWS Glue. Therefore, AWS Glue is the most suitable choice given the constraints and requirements.
pypelyncar 👍 3 Selected: B
EMR provides a managed Hadoop framework that natively supports Apache Pig, Oozie, Spark, and Flink. This allows the company to migrate their existing workloads with minimal code changes, reducing development effort
tgv 👍 2 Selected: B
That's exactly the purpose of EMR. "Amazon EMR is the industry-leading cloud big data solution for petabyte-scale data processing, interactive analytics, and machine learning using open-source frameworks such as Apache Spark, Apache Hive, and Presto." https://aws.amazon.com/emr/
Just_Ninja 👍 3 Selected: A
Glue is Serverless :)
wa212 👍 2 Selected: B
https://docs.aws.amazon.com/ja_jp/emr/latest/ManagementGuide/emr-what-is-emr.html
certplan 👍 2
  • While AWS Glue is a fully managed ETL service and offers serverless capabilities, it might not provide the same level of performance and flexibility as Amazon EMR for handling petabyte-scale workloads with complex processing requirements. - AWS Glue is optimized for data integration, cataloging, and ETL jobs but may not be as well-suited for heavy-duty processing tasks that require frameworks like Apache Spark, Apache Flink, etc., which are commonly used for large-scale data processing. - Documentation on AWS Glue can be found in the AWS Glue Developer Guide https://docs.aws.amazon.com/glue/index.html.
certplan 👍 2
A. AWS Glue: AWS Glue is a fully managed extract, transform, and load (ETL) service provided by Amazon Web Services (AWS). It allows users to prepare and load data for analytics purposes B. Amazon EMR: Amazon Elastic MapReduce (EMR) is a cloud-based big data platform provided by AWS. It allows users to process and analyze large amounts of data using popular frameworks such as Apache Hadoop, Apache Spark, Apache Hive, Apache HBase, and more. https://docs.aws.amazon.com/emr/index.html https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-best-practices.html https://docs.aws.amazon.com/emr/latest/ManagementGuide/emr-manage.html https://docs.aws.amazon.com/emr/latest/DeveloperGuide/emr-developer-guide.html As per the AWS/Amazon docs, option B specifically calls out it out with the specific features/options that the question asked directly about.
GiorgioGss 👍 1 Selected: B
https://docs.aws.amazon.com/emr/latest/ReleaseGuide/emr-release-components.html
TonyStark0122 👍 1
A. AWS Glue
[Removed] 👍 1 Selected: B
https://aws.amazon.com/emr/features/

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Amazon EMR (B) is the industry-leading cloud big data solution for petabyte-scale data processing using open-source frameworks like Apache Spark, Hive, and Presto. Crucially, EMR natively supports Apache Pig, Oozie, HBase, and Flink, allowing the company to migrate their existing workloads with minimal code changes. The question mentions exploring serverless options; Amazon EMR Serverless provides a serverless runtime environment for these applications, eliminating cluster management overhead while maintaining the necessary performance and framework compatibility.

Why the Other Options Are Wrong

AWS Glue (A) is a serverless ETL service, but it uses Apache Spark and PySpark primarily and does not natively support Apache Pig, Oozie, or HBase, making it unsuitable for this specific migration without significant rewrites. Amazon Lambda (C) is limited by execution time (15 minutes max) and memory constraints, making it incapable of processing petabytes of data efficiently. Amazon Redshift (D) is a data warehouse for analytics, not an ETL platform for running distributed computing frameworks like Flink or Spark directly for general-purpose data processing pipelines.

Community Comment Notes

Community consensus strongly favors EMR due to framework compatibility. As milofficial noted, "Glue is like the more good-looking one, but weaker brother of EMR" regarding scale. pypelyncar correctly pointed out that EMR "provides a managed Hadoop framework that natively supports Apache Pig, Oozie, Spark, and Flink." Ell89 confirmed that "Glue doesnt natively support Pig, HBase and Flink," which is the decisive factor here.

Official Reference

Exam Strategy

When a question lists specific open-source frameworks (like Pig, Oozie, HBase), always check if the proposed service supports them natively. If the list includes frameworks unsupported by Glue, EMR is usually the intended answer, even if 'serverless' is mentioned, as EMR Serverless exists.

Frequently Asked Questions

Does Amazon EMR support serverless?

Yes, Amazon EMR Serverless allows you to run Spark, Hive, and other applications without managing clusters, reducing operational overhead.

Why can't I use AWS Glue for Pig and Oozie?

AWS Glue is optimized for Spark and PySpark jobs and does not provide native runtime support for Apache Pig, Oozie, or HBase.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide