Querying Transaction Dates in Centrally Stored S3 CSV Objects with Athena CTAS and No Data Movement
An ML engineer needs to process thousands of existing CSV objects and new CSV objects that are uploaded. The CSV objects are stored in a central Amazon S3 bucket and have the same number of columns. One of the columns is a transaction date. The ML engineer must query the data based on the transaction date. Which solution will meet these requirements with the LEAST operational overhead?
Community Votes
100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
Amazon Athena queries data in place in S3 with standard SQL, so a CTAS statement materializes a table selected by transaction date directly from the central bucket with no data movement, no second bucket, and no pipeline to operate.
Thousands of existing and newly uploaded CSV objects sit in a central S3 bucket with the same column structure, including a transaction date column, and the engineer must query the data by transaction date with the least operational overhead. The data stays in the central bucket and new objects keep arriving.
Setting up a second S3 bucket with replication, or a Spark or Firehose pipeline to copy data before querying. Each of those adds buckets, jobs, and operational surface area, and Firehose in particular cannot consume from S3 at all as a source.
Community Discussion (4 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The requirement is to query centrally stored S3 CSV objects by transaction date with the least operational overhead, and Athena queries data directly where it already lives in S3 using standard SQL, with no data movement or transformation required. A CREATE TABLE AS SELECT statement can select the records filtered by transaction date from the central bucket and materialize the result, so new uploads keep flowing into the same source location without any pipeline. The vote was unanimous at 100 for A. motk123 noted that Athena queries S3 data with SQL without movement or transformation and that CTAS stores the selected results in S3, and feelgoodfactor called it the simplest and most efficient path with minimal operational effort.Why the Other Options Are Wrong
Creating a new bucket with S3 replication and using S3 Object Lambda (B) requires standing up replication and a Lambda-backed access point just to filter by date; S3 Object Lambda is built for on-the-fly object transformation rather than efficient querying, and ninomfr64 judged the approach cumbersome. Creating a new bucket and using AWS Glue for Spark to query and store results (C) does work technically, as ninomfr64 conceded that Spark SQL can query files, but it means provisioning and operating a Spark job and an unnecessary second bucket where Athena needs neither. Using Data Firehose to transfer data to a new bucket and trigger Lambda to query it (D) fails on a hard constraint, because as ninomfr64 pointed out Firehose cannot consume from S3 directly, so it cannot serve as the transfer path from the central bucket.Community Comment Notes
The community was unanimous at 100 for A, with full agreement in the comments. ninomfr64 delivered the sharpest elimination, accepting that options B and C could technically work but calling them cumbersome and more work than Athena, and identifying Firehose's inability to read from S3 as the disqualifier for D. motk123 systematically dismissed the alternatives, noting that S3 Object Lambda is for on-the-fly transformation rather than querying, which is the same distinction that makes Athena the natural fit.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →