How to Speed Up Redshift COPY from an S3 Data Lake?

Answer Correct answer: D — Create a manifest file listing all S3 data file locations and use a single COPY command to load them in parallel into Amazon Redshift.

A company uses Amazon S3 as a data lake. The company sets up a data warehouse by using a multi-node Amazon Redshift cluster. The company organizes the data files in the data lake based on the data source of each data file. The company loads all the data files into one table in the Redshift cluster by using a separate COPY command for each data file location. This approach takes a long time to load all the data files into the table. The company must increase the speed of the data ingestion. The company does not want to increase the cost of the process. Which solution will meet these requirements?

  1. Use a provisioned Amazon EMR cluster to copy all the data files into one folder. Use a COPY command to load the data into Amazon Redshift.
  2. Load all the data files in parallel into Amazon Aurora. Run an AWS Glue job to load the data into Amazon Redshift.
  3. Use an AWS Give job to copy all the data files into one folder. Use a COPY command to load the data into Amazon Redshift.
  4. Create a manifest file that contains the data file locations. Use a COPY command to load the data into Amazon Redshift. Correct Answer

Community Votes

D
100%

100% of anonymous learners picked answer D. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The question tests the Redshift best practice of one COPY command with a manifest for multiple S3 files; the trap is adding EMR, Glue, or Aurora, which adds cost and an unnecessary intermediate copy step.

Amazon Redshift loads data fastest from Amazon S3 when a single COPY command reads all files through a manifest instead of issuing one COPY command per file location. A manifest lists every data file, letting Redshift parallelize the load across slices without introducing extra services or cost.

Choosing the AWS Glue option because it sounds like a managed way to consolidate files, but running a Glue job adds cost and still requires an extra copy before the Redshift COPY, so it does not meet the no-cost increase requirement.

Community Discussion (3 comments)

andrologin 👍 1 Selected: D
D is the right answer based on the docs in this page https://docs.aws.amazon.com/redshift/latest/dg/c_best-practices-single-copy-command.html
HunkyBunky 👍 4 Selected: D
Only D makes sense https://docs.aws.amazon.com/redshift/latest/dg/c_best-practices-single-copy-command.html
Bmaster 👍 1
D is good https://docs.aws.amazon.com/redshift/latest/dg/c_best-practices-single-copy-command.html https://docs.aws.amazon.com/redshift/latest/dg/loading-data-files-using-manifest.html

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

A Redshift COPY command can load multiple Amazon S3 objects in one operation when you provide a manifest that explicitly lists each file location. This directly addresses the slow pattern of running a separate COPY command for every source prefix, because Redshift can distribute and parallelize the load across all slices of the multi-node cluster. The manifest preserves the existing data-lake organization by source while avoiding the overhead of repeated COPY command startup and commit cycles. AWS documentation explicitly recommends using a single COPY command rather than multiple COPY commands for this scenario. No additional cluster, ETL job, or database service is required, so ingestion gets faster without increasing cost.

Why the Other Options Are Wrong

Option A uses a provisioned Amazon EMR cluster to copy files into one folder, which adds EMR cost and an extra copy step before the Redshift COPY. Option B introduces Amazon Aurora and an AWS Glue job, creating cross-service data movement and cost that the requirement forbids. Option C uses an AWS Glue job, likely a typo for Glue, to consolidate files, but that job costs money, takes time, and still leaves a single COPY after an intermediate copy. Only the manifest approach in option D lets one COPY command target all S3 file locations directly and in parallel without extra infrastructure.

Community Comment Notes

HunkyBunky summarized the community view by saying "Only D makes sense" and linked the Redshift best-practices page for a single COPY command. Andrologin also endorsed D based on that same AWS documentation. Bmaster agreed that D is good and pointed to both the single-COPY best-practice page and the Redshift manifest documentation, reinforcing that the manifest is the intended mechanism for loading many files efficiently.

Official Reference

Exam Strategy

When a Redshift COPY from S3 is slow and cost must stay flat, look for the option that keeps the load inside Redshift rather than adding EMR, Glue, or Aurora. Manifest files are the standard exam signal for loading multiple S3 objects with a single parallel COPY command.

Frequently Asked Questions

Why is a single COPY command with a manifest faster than separate COPY commands?

Redshift can parallelize one COPY across all slices and read every manifest-listed S3 object in a single operation, avoiding repeated command startup and commit overhead.

Why not use an AWS Glue job to consolidate the data files first?

A Glue job adds cost and an intermediate copy step, while the manifest lets Redshift load the existing S3 files directly without extra services.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide