Redshift Data Sharing for Cross-Cluster Access

Answer Correct answer: A — Set up the sales team BI cluster as a consumer of the ETL cluster by using Redshift data sharing.

A company maintains an Amazon Redshift provisioned cluster that the company uses for extract, transform, and load (ETL) operations to support critical analysis tasks. A sales team within the company maintains a Redshift cluster that the sales team uses for business intelligence (BI) tasks. The sales team recently requested access to the data that is in the ETL Redshift cluster so the team can perform weekly summary analysis tasks. The sales team needs to join data from the ETL cluster with data that is in the sales team's BI cluster. The company needs a solution that will share the ETL cluster data with the sales team without interrupting the critical analysis tasks. The solution must minimize usage of the computing resources of the ETL cluster. Which solution will meet these requirements?

  1. Set up the sales team BI cluster as a consumer of the ETL cluster by using Redshift data sharing. Correct Answer
  2. Create materialized views based on the sales team's requirements. Grant the sales team direct access to the ETL cluster.
  3. Create database views based on the sales team's requirements. Grant the sales team direct access to the ETL cluster.
  4. Unload a copy of the data from the ETL cluster to an Amazon S3 bucket every week. Create an Amazon Redshift Spectrum table based on the content of the ETL cluster.

Community Votes

A
70%
D
30%

70% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The core concept is minimizing resource usage on the source (producer) cluster while enabling cross-cluster queries, which Redshift Data Sharing achieves by decoupling compute from storage access.

This question tests the use of Amazon Redshift Data Sharing to securely share live data between clusters without impacting producer performance. The correct solution leverages this feature to allow the sales team to join ETL data with their BI cluster data efficiently.

Candidates often choose Option D (S3/Spectrum) because they focus on the 'weekly' frequency or misinterpret Spectrum as having zero impact, failing to realize that Data Sharing is specifically designed for this low-overhead cross-cluster scenario.

Community Discussion (12 comments)

arvehisa 👍 5 Selected: A
A: redshift data sharing: https://docs.aws.amazon.com/redshift/latest/dg/data_sharing_intro.html With data sharing, you can securely and easily share live data across Amazon Redshift clusters. B: materialized view is only within 1 redshift cluster, across different tables
lucas_rfsb 👍 5 Selected: D
In my opinion using Redshift Data Sharing will consume less resources. 'D' envolves using a S3 bucket.
motk123 👍 2
Seems that the performance of the critical ETL cluster should not be affected when using data sharing, so the answer is likely A: https://docs.aws.amazon.com/redshift/latest/dg/data_sharing_intro.html Supporting different kinds of business-critical workloads – Use a central extract, transform, and load (ETL) cluster that shares data with multiple business intelligence (BI) or analytic clusters. This approach provides read workload isolation and chargeback for individual workloads. You can size and scale your individual workload compute according to the workload-specific requirements of price and performance. https://docs.aws.amazon.com/redshift/latest/dg/considerations.html The performance of the queries on shared data depends on the compute capacity of the consumer clusters.
wimalik 👍 1
A as Redshift data sharing allows you to share live data across Redshift clusters without having to duplicate the data. This feature enables the sales team to access the data from the ETL cluster directly without interrupting the critical analysis tasks or overloading the ETL cluster's resources. The sales team can join this shared data with their own data in the BI cluster efficiently.
San_Juan 👍 1 Selected: D
"The solution must minimize usage of the computing resources of the ETL cluster." That is key. You shouldn't use ETL cluster, so unload data to S3 and run queries in a separate Redshift Spectrum database. ETL cluster do nothing meanwhile.
VerRi 👍 3 Selected: A
Typetical Redshift data sharing use case
valuedate 👍 2
key words: "weekly" "The solution must minimize usage of the computing resources of the ETL cluster." Answer:D
d8945a1 👍 4 Selected: A
Typical usecase of datasharing in Redshift. The question mentions that - 'team needs to join data from the ETL cluster with data that is in the sales team's BI cluster.' This is possible with datashare.
jasango 👍 3 Selected: D
The spectrum table is accessed from the sales cluster with zero impact on the ETL cluster.
certplan 👍 1
Options A, B, and C involve granting the sales team direct access to the ETL cluster, which could potentially impact the performance of the ETL cluster and interfere with its critical analysis tasks. Option D provides a more isolated and scalable approach by leveraging Amazon S3 and Redshift Spectrum for data sharing while minimizing the usage of the ETL cluster's computing resources. https://docs.aws.amazon.com/redshift/latest/dg/c-using-spectrum-sharing-data.html https://docs.aws.amazon.com/redshift/latest/dg/c_best-practices-design-tables.html
GiorgioGss 👍 5 Selected: A
Initially I would go with B but that definitely will use more resource.
[Removed] 👍 4 Selected: A
To share data between Redshift clusters and meet the requirements of sharing ETL cluster data with the sales team without interrupting critical analysis tasks and minimizing the usage of the ETL cluster's computing resources, Redshift Data Sharing is the way to go https://docs.aws.amazon.com/redshift/latest/dg/data_sharing_intro.html "Supporting different kinds of business-critical workloads – Use a central extract, transform, and load (ETL) cluster that shares data with multiple business intelligence (BI) or analytic clusters. This approach provides read workload isolation and chargeback for individual workloads. You can size and scale your individual workload compute according to the workload-specific requirements of price and performance"

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option A is the correct choice because Amazon Redshift Data Sharing allows a producer cluster (ETL) to share live data with consumer clusters (BI) without duplicating the data. This mechanism ensures that queries run on the consumer cluster's resources, thereby minimizing the impact on the ETL cluster's computing power and preventing interruptions to critical analysis tasks.

Why the Other Options Are Wrong

Options B and C involve granting direct access to the ETL cluster. This would cause the sales team's queries to compete for CPU and memory with the ETL workloads, violating the requirement to minimize resource usage and avoid interruption. Option D involves unloading data to S3 and using Redshift Spectrum. While Spectrum queries also use consumer cluster resources, it introduces latency due to data movement and storage management, making Data Sharing the more native and efficient solution for inter-cluster sharing.

Community Comment Notes

The community largely agrees with Option A, citing official documentation that describes Data Sharing as ideal for separating ETL and BI workloads. Some users like Lucas_rfsb and jasango argue for Option D, claiming Spectrum has zero impact, but they overlook that Data Sharing is the specific architectural pattern for this problem. As arvehisa noted, "redshift data sharing" is the key term here. One user correctly pointed out that materialized views are limited to a single cluster context, ruling out B.

Official Reference

Exam Strategy

When a question mentions sharing data between clusters while protecting the source cluster's performance, look for "Data Sharing" or "Data Exchange" rather than replication or direct access. Always prioritize solutions that keep compute isolated to the querying side.

Frequently Asked Questions

Why not use Redshift Spectrum instead of Data Sharing?

Spectrum requires data in S3 and is best for ad-hoc queries on large datasets. Data Sharing is optimized for live, structured data exchange between clusters with lower overhead.

Does Data Sharing copy the data to the consumer cluster?

No, it does not duplicate the data. It provides a pointer to the producer's data, allowing the consumer to query it directly using its own compute resources.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide