Enabling Apache Spark in Amazon Athena
A company uses Amazon Athena to run SQL queries for extract, transform, and load (ETL) tasks by using Create Table As Select (CTAS). The company must use Apache Spark instead of SQL to generate analytics. Which solution will give the company the ability to use Spark to access Athena?
Community Votes
72% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests the specific architectural component required to switch Athena's execution engine from SQL to Spark; the trap is confusing the data source connector with the execution environment (workgroup).
To use Apache Spark instead of SQL for ETL tasks in Amazon Athena, you must create an Athena workgroup configured with the Spark engine. This configuration enables notebook-based analytics within the Athena environment.
Candidates often select 'Athena data source' because it conceptually bridges data access, but this is incorrect for enabling the Spark engine itself within Athena's managed service.
Community Discussion (10 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Amazon Athena allows users to run queries using either Presto SQL or Apache Spark. To utilize Apache Spark, you must explicitly configure an Athena workgroup to use the Spark engine. According to AWS documentation, "To use Apache Spark in Amazon Athena, you create an Amazon Athena workgroup that uses a Spark engine." Once created, this workgroup provides the necessary context and resources to open notebooks and execute Spark code.Why the Other Options Are Wrong
Athena query settings (A) manage individual query parameters like output location but do not change the underlying engine. An Athena data source (C) typically refers to connectors or metadata catalogs that allow external tools (like Spark running on EMR) to read Athena tables, but it does not enable Spark inside Athena. The Athena query editor (D) is an interface for writing SQL queries, not a configuration for changing the processing engine.Community Comment Notes
The community consensus strongly favors option B, supported by direct references to the official AWS Getting Started guide. Several commenters noted that while data sources facilitate interaction between Spark and Athena data, the prerequisite step to actually use Spark within Athena is creating a Spark-enabled workgroup. As one commenter stated, "You need an Athena workgroup as a prerequisite to use Apache Spark."Official Reference
Exam Strategy
When a question asks how to 'use' a specific engine or feature within a managed AWS service, look for the configuration object (like a Workgroup, Cluster, or Instance Group) that activates that feature. Do not confuse data connectivity options with execution environment configurations.
Frequently Asked Questions
Why isn't 'Athena data source' the correct answer?
An Athena data source allows external Spark applications (e.g., on EMR) to read Athena tables. It does not enable the Spark engine inside Athena itself.
Can I use Spark without a workgroup?
No. You must first create a workgroup specifically configured with the Spark engine before you can open notebooks or run Spark code in Athena.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →