Athena Federated Query for Multi-Source Iceberg Pipelines

Answer Correct answer: B — Use Amazon Athena federated query connectors to read, join, and merge data from Redshift, Teradata, and BigQuery into an Iceberg table.

A company has three subsidiaries. Each subsidiary uses a different data warehousing solution. The first subsidiary hosts its data warehouse in Amazon Redshift. The second subsidiary uses Teradata Vantage on AWS. The third subsidiary uses Google BigQuery. The company wants to aggregate all the data into a central Amazon S3 data lake. The company wants to use Apache Iceberg as the table format. A data engineer needs to build a new pipeline to connect to all the data sources, run transformations by using each source engine, join the data, and write the data to Iceberg. Which solution will meet these requirements with the LEAST operational effort?

  1. Use native Amazon Redshift, Teradata, and BigQuery connectors to build the pipeline in AWS Glue. Use native AWS Glue transforms to join the data. Run a Merge operation on the data lake Iceberg table.
  2. Use the Amazon Athena federated query connectors for Amazon Redshift, Teradata, and BigQuery to build the pipeline in Athena. Write a SQL query to read from all the data sources, join the data, and run a Merge operation on the data lake Iceberg table. Correct Answer
  3. Use the native Amazon Redshift connector, the Java Database Connectivity (JDBC) connector for Teradata, and the open source Apache Spark BigQuery connector to build the pipeline in Amazon EMR. Write code in PySpark to join the data. Run a Merge operation on the data lake Iceberg table.
  4. Use the native Amazon Redshift, Teradata, and BigQuery connectors in Amazon Appflow to write data to Amazon S3 and AWS Glue Data Catalog. Use Amazon Athena to join the data. Run a Merge operation on the data lake Iceberg table.

Community Votes

B
54%
A
46%

54% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The exam tests knowledge of Athena's federated query capabilities and its native support for Apache Iceberg. The common trap is assuming that 'pipeline' implies using Glue or EMR, whereas managed SQL-based solutions often offer less operational effort for specific use cases.

This question evaluates the ability to select AWS services for aggregating heterogeneous data sources into an S3 data lake with minimal operational overhead. It establishes that Athena Federated Query is the optimal solution for connecting to and querying external data warehouses directly without building complex ETL infrastructure.

Candidates frequently choose Option A (AWS Glue) because it is a dedicated ETL service. However, maintaining native connectors for three distinct engines and managing serverless Glue jobs involves more operational configuration than using Athena's managed connectors.

Community Discussion (9 comments)

bad1ccc 👍 1 Selected: B
https://docs.aws.amazon.com/athena/latest/ug/federated-queries.html
Palee 👍 1 Selected: D
The requirement is to aggregate the data in S3. Only option has exclusively called this out. So Ans D is correct
MerryLew 👍 1 Selected: A
Athena can be used to build certain types of data pipelines, particularly when the primary focus is on ad-hoc analysis and querying large datasets stored in S3 without the need for complex data transformations, but for more intricate data processing and heavy ETL operations, other AWS services like Glue are often more suitable due to their dedicated data processing capabilities.
Eeshav15 👍 1 Selected: A
Glue is the right tool to build pipeline
michele_scar 👍 2 Selected: B
https://docs.aws.amazon.com/athena/latest/ug/connectors-available.html
Eleftheriia 👍 3 Selected: B
Would it be B "If you have data in sources other than Amazon S3, you can use Athena Federated Query to query the data in place or build pipelines that extract data from multiple data sources and store them in Amazon S3. With Athena Federated Query, you can run SQL queries across data stored in relational, non-relational, object, and custom data sources." https://docs.aws.amazon.com/athena/latest/ug/connect-to-a-data-source.html
kupo777 👍 3
Correct Answer: B Use the Amazon Athena federated query connectors for Amazon Redshift, Teradata, and BigQuery to build the pipeline in Athena. Write a SQL query to read from all the data sources, join the data, and run a Merge operation on the data lake Iceberg table.
ae35a02 👍 3 Selected: A
AWS GLUE has native connectors to Redshift, BigQuery and Terradata, and integrates with Iceberg format. Athena is not for building Pipelines, AppFlow is for transfering data from Saas applications
Parandhaman_Margan 👍 2
Answer:A

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option B is correct because Amazon Athena Federated Query allows you to query data in external data sources like Redshift, Teradata, and BigQuery directly. The documentation confirms that "If you have data in sources other than Amazon S3, you can use Athena Federated Query to query the data in place or build pipelines that extract data from multiple data sources and store them in Amazon S3." Since the goal is to join data and write to Iceberg with the LEAST operational effort, leveraging SQL queries against managed connectors is superior to writing custom code or managing clusters.

Why the Other Options Are Wrong

Option A (AWS Glue) requires configuring and maintaining specific connectors for each engine and writing Glue Python shell scripts or Spark jobs, which increases operational effort compared to pure SQL. Option C (Amazon EMR) requires managing a cluster (or EMR Serverless), installing drivers, and writing PySpark code, representing the highest operational overhead. Option D (AppFlow) is designed primarily for SaaS-to-AWS data transfer and does not support direct integration with these specific data warehouse engines as connectors for this type of analytical pipeline.

Community Comment Notes

Many learners were confused by the word "pipeline," with several arguing that "Glue is the right tool to build pipeline" or noting that "Athena is not for building Pipelines." However, the official AWS documentation explicitly supports using Athena for pipelines that extract data from multiple sources. One commenter noted that Athena Federated Query can "run SQL queries across data stored in relational, non-relational, object, and custom data sources," validating the approach.

Official Reference

Exam Strategy

When asked for the 'LEAST operational effort', always prioritize fully managed, serverless services over those requiring code or cluster management. If the task involves joining data from disparate sources, check if a managed connector exists before defaulting to ETL tools like Glue or EMR.

Frequently Asked Questions

Is Athena considered a valid tool for building data pipelines?

Yes. AWS documentation states Athena Federated Query can be used to build pipelines that extract data from multiple sources and store them in Amazon S3.

Why not use AWS Glue for this scenario?

While Glue is a strong ETL tool, it requires more setup for diverse connectors and scripting compared to Athena's SQL-based federated queries, violating the 'least operational effort' constraint.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide