DEA-C01 — Frequently Asked Questions
Community-vetted answers to 84 common questions about this exam.
Questions from real practice questions
Each Q&A comes from a specific community question — follow the link for its full analysis.
How to Fix an AWS Glue S3 VPC Gateway Endpoint Error?
An S3 gateway endpoint is attached to route tables; the Glue job's subnet needs a route to the endpoint's S3 prefix list so packets can reach S3 without a NAT or internet gateway.
Glue serverless jobs do not use a security group to reach S3 through a gateway endpoint; gateway endpoints rely on route tables, so inbound security group rules do not apply.
Cost-Effective Querying of Compressed Data for Audits
S3 Select reduces data scanned during queries, but storing large volumes of historical audit data in S3 Standard incurs high monthly storage fees compared to Glacier.
No, Athena operates on data in S3 Standard or similar tiers. It cannot directly query objects in Glacier without first restoring them, which defeats the purpose of cost optimization.
How Do You Schedule a Redshift Stored Procedure Daily Cost-Effectively?
Lambda needs custom code, IAM permissions, and VPC or Data API connectivity to reach Redshift, adding operational overhead even if invocation cost is low.
Yes, scheduled queries in query editor v2 run SQL against your Redshift cluster and can include CALL statements to invoke stored procedures.
How to Fix EventBridge Lambda AccessDeniedException?
The trust policy only lets the Lambda service assume the execution role to run code. EventBridge invokes Lambda through the function's resource-based policy, not that trust policy.
Option B covers both sides: the EventBridge rule needs invoke permissions, and the Lambda resource-based policy must allow the EventBridge principal. Missing either can produce AccessDeniedException.
How to Apply Two Layers of Server-Side Encryption to S3 Uploads?
The S3 Encryption Client encrypts data on the client side before upload, so only the SSE-KMS portion is server-side; the requirement demands two server-side encryption layers.
Yes. DSSE-KMS is applied by Amazon S3 at upload time, so the Lambda function can use a normal PutObject request and S3 transparently applies both server-side encryption layers.
AWS Glue Job Bookmarks Reprocessing Cause
Missing ACLs typically cause access denied errors, not silent reprocessing. Bookmark state persistence requires an explicit commit.
No. Concurrency controls parallel task execution but does not influence how bookmark checkpoints are saved or retrieved.
Orchestrate Python and Bash Scripts with Least Overhead
Step Functions requires rewriting workflows into JSON state machines, forcing code refactoring. MWAA supports existing Python/Bash DAGs directly.
No, MWAA is a managed service where AWS handles the infrastructure, allowing you to focus only on your workflow code.
Cost-Effective ACID Querying on S3 Data
Yes, Athena supports ACID transactions for Hive transactional tables stored in S3, ensuring data integrity.
While serverless, Athena is more cost-effective for ad-hoc S3 queries and offers native ACID support without needing a Redshift cluster context.
AWS DMS Replication Instance Region for Cross-Region Migration
DMS instances are regional resources. For Redshift, the instance MUST be in the same region as the target cluster due to VPC and endpoint constraints.
No. While often placed near the source for performance, it must be in the same region as the target if the target is Amazon Redshift.
Kinesis Data Firehose CSV to Parquet Conversion
Firehose natively converts JSON to Parquet/ORC. For CSV, it typically requires transformation to JSON first, but exam logic prioritizes built-in features over Lambda for 'least effort'.
Lambda requires custom code development and maintenance. The question asks for the LEAST development effort, making built-in Firehose features preferable if applicable.
Low Latency Real-Time Sensor Dashboard with Kinesis and Timestream
QuickSight is designed for BI and typically has refresh intervals ranging from minutes to seconds, whereas Grafana can query time-series databases like Timestream for sub-second real-time updates.
No, Kinesis Data Firehose does not support Amazon Timestream as a native destination. You must use a consumer application like Apache Flink or Kinesis Client Library to ingest and write to Timestream.
Real-Time Network Outage Detection with Kinesis and Flink
Activity Streams capture SQL operations for security auditing, not business metrics like network usage volumes.
No, polling every minute introduces significant delay compared to the near-instant processing of a Flink stream.
Minimizing SQS Data Loss During Downtime
No. Visibility timeout only affects when a message becomes available after being received, not how long it stays in the queue before expiration.
DLQs do not prevent expiration, but they capture messages that fail processing repeatedly, preventing permanent deletion and allowing recovery.
Reclaiming Storage in Amazon Redshift Materialized Views
DELETE only marks rows as deleted. The actual space is reclaimed when VACUUM runs, making DELETE slow for large-scale removal.
Generally no, TRUNCATE applies to base tables. This question likely assumes a context where the underlying table is truncated or refers to streaming ingestion specifics.
Least Overhead Ingestion to OpenSearch via Kinesis
Logstash is open-source and typically requires managing EC2 instances or clusters to run, whereas Kinesis Data Firehose is a fully managed service that handles scaling and maintenance automatically.
You can use Lambda to process Kinesis streams, but integrating it directly with OpenSearch often requires custom code for batching and retry logic. Firehose provides built-in buffering, compression, and retry mechanisms, reducing overhead.
Amazon Redshift Third-Party IdP Authentication Setup
Yes, it refers to executing the CREATE IDENTITY_PROVIDER SQL command against the Redshift cluster, rather than just clicking buttons in the console without backend config.
While you must configure the IdP to trust Redshift first, the question asks what the data engineer should do to set up the mechanism in Redshift. The exam key prioritizes the Redshift-side registration step.
Git Rebase vs Merge Before Pull Request
Git pull creates a merge commit, which can clutter the history with non-linear merge nodes. Rebasing moves your commits to the tip of master, keeping history linear.
Rebasing rewrites commit history. If others have cloned your feature branch, their local copies will diverge, requiring force pushes and causing collaboration issues.
Why Does QuickSight Report Insufficient Permissions for Athena S3 Data?
A missing QuickSight service role causes an authentication or authorization failure before Athena queries run, not the specific Athena insufficient-permissions error; QuickSight normally auto-creates aws-quicksight-service-role-v0.
If the S3 objects or Athena query results are encrypted with SSE-KMS, QuickSight's role must have kms:Decrypt and the key policy must allow it, otherwise queries fail with insufficient permissions.
How to Query S3, RDS, DynamoDB, and Redshift with SQL?
Redshift Spectrum only queries external tables in Amazon S3 through the Glue Data Catalog; it cannot directly query DynamoDB tables or RDS SQL Server. You would need extra federation or export steps, so option B adds overhead.
Option C adds Glue jobs to convert JSON into Parquet or CSV, which is unnecessary ETL, and it still only queries S3. It does not provide a unified query path to RDS, DynamoDB, or Redshift.
How Do You Fix SageMaker Studio Access Denied for Glue Sessions?
It grants broad SageMaker permissions but not the sts:AssumeRole call a Studio user needs before AWS Glue interactive sessions can run, so AccessDenied can persist.
No. AddAssociation is an Amazon SageMaker API operation, not an STS action; cross-service access is granted with sts:AssumeRole, which option B specifies.
Most Cost-Effective AWS Glue ETL for Small Daily CSV Files?
The shell job can run on 1/16 DPU and pandas handles sub-100 MB CSVs in memory, while a PySpark job bills full DPUs plus Spark startup overhead the small files never need.
When files or volumes are large enough that data cannot fit in a single node's memory or the transform must be distributed across a Spark cluster; that scalability is wasted on under-100 MB daily CSVs.
Which AWS service most cost-effectively orchestrates an Amazon EMR ETL pipeline?
Glue Workflows only coordinates AWS Glue jobs and crawlers, so it cannot directly manage EMR cluster steps or non-Glue AWS services. The question specifically requires orchestrating an existing EMR ETL pipeline.
MWAA runs an always-on managed Airflow environment with worker and scheduler charges, while Step Functions charges per state transition and scales to zero when idle.
How to Auto-Start a Glue Workflow After S3 File Gateway Transfers?
EventBridge can target an AWS Glue workflow directly, so a Lambda function adds code, permissions, and monitoring overhead without improving the event-driven trigger.
Yes, S3 File Gateway publishes successful upload events to EventBridge, allowing rules to react when a file transfer finishes.
How Do You Combine Live, Current, and Archived Redshift Data?
Spectrum only reads S3 Standard, S3 Intelligent-Tiering, and Glacier Instant Retrieval, so Glacier Flexible Retrieval objects must be restored and copied to a supported class before Spectrum can query them.
No. Federated Query adds direct read access to live Aurora PostgreSQL for analytics, but the current Redshift data and the S3 history still come from the existing ETL run and the monthly UNLOAD job.
Most Cost-Effective VPC Flow Log Analytics Solution?
Parquet is columnar and compressed, so Athena reads fewer bytes per query and S3 stores less data. Because Athena charges by data scanned, this lowers analytics cost.
Yes. When you create a flow log with an S3 destination, you can choose Parquet as the output format, so no separate conversion service is required.
Which Kinesis Option Delivers IoT Sensor Data to S3 with Least Latency?
Kinesis Data Streams stores records for consumers but has no native S3 sink. You need a KCL application or Firehose to read records and write them to S3.
Firehose historically enforced a 60-second minimum, and AWS later added zero buffering; either way, adding Managed Flink before Firehose introduces more processing latency than a direct KCL write.
How to Trigger a Lambda CSV-to-Parquet Conversion on S3 Upload?
S3 event notifications can invoke a Lambda function directly by using the function ARN as the destination. Adding SNS only introduces an extra topic, subscription and permissions to manage for no functional gain.
No. S3 supports specific event families such as s3:ObjectCreated: and s3:ObjectRemoved:; s3:* cannot be configured, which is why option C would never trigger the Lambda conversion.
Most Cost-Effective Athena and QuickSight SPICE for S3 Clickstream?
Direct SQL makes each dashboard interaction query Athena, scanning S3 data repeatedly and increasing cost and latency. SPICE caches the dataset once and serves users from memory with a daily refresh.
Redshift requires a provisioned cluster and ongoing management, which is not the most cost-effective option when the data already sits in S3. Athena queries S3 serverlessly and scales with daily updates.
Which Service Supports Hybrid On-Premises and Cloud Orchestration?
AWS Glue is a serverless AWS-native ETL and catalog service; it is not an open-source workflow engine that can run on-premises with the same definitions.
Yes, MWAA is managed Apache Airflow, so the same DAG definitions and plugins can generally run on-premises and in AWS, preserving orchestration portability.
Which AWS NoSQL Database Meets Global OLTP with Least Overhead?
Keyspaces targets Cassandra-compatible workloads and can require more capacity and migration design; the question asks for the least-overhead global OLTP solution, which DynamoDB addresses natively.
DocumentDB can deliver low-latency document reads, but it is optimized for MongoDB compatibility rather than the global, serverless OLTP profile described in the question.
How to Prevent Amazon Athena Queries From Queueing?
The result limit only caps output size or bytes scanned; queueing is driven by whether the workgroup has available query-processing capacity, which option A does not change.
No. Adding users changes who can submit queries, not how much capacity the workgroup has; it can increase contention and make queueing worse.
How to Make Amazon Athena Queries on CSV Files Run Faster?
Compression only shrinks bytes on disk; CSV is still row-based, so Athena reads every column for a single-column filter. Parquet adds column pruning plus predicate pushdown, cutting scanned data far more.
JSON is row-oriented text and usually larger than CSV, so it keeps the same full-row scan pattern. Only the columnar layout of Parquet lets Athena read just the selected column.
Least-Effort Near Real-Time MySQL to Redshift Integration
A scheduled Glue job runs in batches, so it cannot continuously capture PLM transaction changes in near real time, and it requires writing and maintaining ETL logic, which is more development effort than an AWS DMS CDC task.
Yes. AWS DMS supports MySQL as a source and Amazon Redshift as a target, and a full load plus CDC task can apply ongoing changes to Redshift continuously.
How to Speed Up Redshift COPY from an S3 Data Lake?
Redshift can parallelize one COPY across all slices and read every manifest-listed S3 object in a single operation, avoiding repeated command startup and commit overhead.
A Glue job adds cost and an intermediate copy step, while the manifest lets Redshift load the existing S3 files directly without extra services.
How to Enforce TLS 1.2 on an AWS Transfer Family Server?
A certificate supplies the identity and key material for the endpoint, but it does not restrict which protocol versions the server accepts. Only the Transfer Family security policy sets the minimum TLS version.
No. Security groups filter traffic at Layers 3 and 4 using IPs, ports, and protocols such as TCP, so they cannot inspect or limit the TLS version negotiated inside a connection.
Which AWS Glue Feature Enables Incremental ETL Ingestion?
Triggers only start a crawler or job on a schedule or event; they do not remember which S3 objects were already processed, so each run would reprocess the entire dataset without job bookmarks.
Yes. Bookmark state is tracked per transformation regardless of file compression, so gzip or bzip2 objects already read are skipped and only new or changed files are ingested.
How Do You Achieve Exactly-Once Delivery in a Kinesis Data Streams Pipeline?
A timed-out PutRecord may have succeeded, so the application must retry and Kinesis stores both copies. Duplicate ingestion is inherent to at-least-once delivery and cannot be prevented at the source.
Checkpoints guarantee exactly-once state consistency inside the Flink job, but they cannot remove duplicates that PutRecord retries already wrote into the Kinesis stream before the job read them.
Least-Effort Way to Count Distinct Customers from an .xls File in S3?
Athena does not natively read .xls; you must convert the file to a supported format and catalog it, adding operational steps compared with DataBrew.
Yes. DataBrew supports Excel files stored in Amazon S3 and provides aggregate recipe functions such as COUNT_DISTINCT for distinct customer counts.
How to stream Kinesis Data Streams to Redshift Serverless with least overhead?
Firehose delivers to Redshift only through S3 and COPY, adding buffering, extra storage, and pipeline management. Redshift streaming ingestion removes those intermediate steps.
Yes. Streaming ingestion populates materialized views for near real-time data, while the same Redshift Serverless warehouse stores and queries historical tables for yesterday's data.
Which AWS solution detects changing schemas and loads data to S3 with least overhead?
Glue is serverless, so it infers changing schemas with crawlers and runs Apache Spark ETL without clusters to provision or tune, which directly meets the least-operational-overhead requirement.
No. Lambda has a 15-minute maximum runtime and limited packaging, making it impractical for a 1 TB daily pipeline with schema detection across SAP HANA, Kafka, and DynamoDB.
How to Dynamically Redact S3 PII for Multiple Applications?
Bucket policies control access but do not redact PII inside objects, and multiple copies add storage, sync, and governance overhead that the least-overhead requirement rules out.
Yes. It runs Lambda code on S3 GET requests, so you can identify the caller and return redacted data dynamically without modifying the stored dataset.
How to Separate Athena Query Access and Query History?
IAM roles control which AWS API actions a principal can call, but they do not partition Athena's workgroup-level query history or enforce per-workgroup query limits.
Tags let an IAM policy scope Allow or Deny statements to specific workgroups, so users in one account can only run queries and view history for their own use case.
Ready to practice?
Access 100 DEA-C01 questions with instant feedback and detailed explanations.
View DEA-C01 Practice Questions →← Back to DEA-C01 AWS Certified Data Engineer - Associate Study Guide