Least Overhead PII Redaction in S3
A company stores customer data in an Amazon S3 bucket. Multiple teams in the company want to use the customer data for downstream analysis. The company needs to ensure that the teams do not have access to personally identifiable information (PII) about the customers. Which solution will meet this requirement with LEAST operational overhead?
Community Votes
100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The core trap is confusing data discovery with data protection; learners must recognize that Macie only detects PII, whereas S3 Object Lambda actively transforms and redacts it during retrieval.
This question addresses the best practice for protecting Personally Identifiable Information (PII) stored in Amazon S3 with minimal operational overhead, identifying S3 Object Lambda as the optimal solution. It establishes that while detection is one step, automated redaction upon access is required to meet the specific security requirement.
Many users select Amazon Macie because it is the primary service for discovering sensitive data in S3, failing to realize that Macie cannot automatically remove or mask the data from the bucket.
Community Discussion (6 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
S3 Object Lambda is the correct choice because it allows you to add custom code to process data as it is retrieved from S3 without changing your application architecture. By attaching a Lambda function (powered by services like Amazon Comprehend or simple logic), you can detect and redact PII on-the-fly before the data reaches the consumer team. This approach ensures that the raw data remains intact in the source but never exposes PII to downstream analysts, fulfilling the security requirement with zero additional storage or pipeline management overhead.Why the Other Options Are Wrong
Amazon Macie (Option A) is designed for discovery and classification of sensitive data; it alerts you to the presence of PII but does not have the capability to automatically delete or redact the data from the object itself. Option C involves Amazon Data Firehose, which requires setting up a streaming pipeline, increasing operational overhead significantly compared to direct access transformation. Option D suggests using AWS Glue DataBrew to create a separate copy of non-PII data, which doubles storage costs and introduces manual ETL maintenance, violating the 'least operational overhead' constraint.Community Comment Notes
The community consensus strongly supports Option B, with many users noting that Macie is limited to detection. As user paali noted, "Macie will only detect sensitive data, it can't redact it," highlighting the critical distinction between finding data and protecting it. User michele_scar provided a direct link to the AWS tutorial for S3 Object Lambda PII redaction, confirming this is an official AWS recommended pattern for this specific scenario.Official Reference
Exam Strategy
When asked about 'least operational overhead' for data protection, look for serverless, integrated solutions that modify data at the point of use rather than moving data to new pipelines or buckets. Always distinguish between 'discovery' services (like Macie) and 'transformation/protection' services (like Object Lambda).
Frequently Asked Questions
Why isn't Amazon Macie the right answer for removing PII?
Macie discovers and classifies sensitive data but does not have the functionality to automatically redact or remove that data from the S3 objects.
What is the operational overhead difference between S3 Object Lambda and DataBrew?
Object Lambda processes data on-demand with no infrastructure to manage, while DataBrew requires configuring jobs, scheduling, and managing output storage.
Related Analysis
Practice All DEA-C01 Questions
Access 100 questions with complete answers and detailed explanations.
View Full DEA-C01 Practice Test →