Least Overhead PII Redaction in S3

Answer Correct answer: B — Use S3 Object Lambda to access the data, and use Amazon Comprehend to detect and remove PII.

A company stores customer data in an Amazon S3 bucket. Multiple teams in the company want to use the customer data for downstream analysis. The company needs to ensure that the teams do not have access to personally identifiable information (PII) about the customers. Which solution will meet this requirement with LEAST operational overhead?

  1. Use Amazon Macie to create and run a sensitive data discovery job to detect and remove PII.
  2. Use S3 Object Lambda to access the data, and use Amazon Comprehend to detect and remove PII. Correct Answer
  3. Use Amazon Data Firehose and Amazon Comprehend to detect and remove PII.
  4. Use an AWS Glue DataBrew job to store the PII data in a second S3 bucket. Perform analysis on the data that remains in the original S3 bucket.

Community Votes

B
100%

100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The core trap is confusing data discovery with data protection; learners must recognize that Macie only detects PII, whereas S3 Object Lambda actively transforms and redacts it during retrieval.

This question addresses the best practice for protecting Personally Identifiable Information (PII) stored in Amazon S3 with minimal operational overhead, identifying S3 Object Lambda as the optimal solution. It establishes that while detection is one step, automated redaction upon access is required to meet the specific security requirement.

Many users select Amazon Macie because it is the primary service for discovering sensitive data in S3, failing to realize that Macie cannot automatically remove or mask the data from the bucket.

Community Discussion (6 comments)

HagarTheHorrible 👍 3 Selected: B
it is not A, Macie can only detect the PII
WarPig666 👍 2 Selected: B
A can’t be correct. Macie can discover PII, but not automatically redact it.
paali 👍 2 Selected: B
Macie will only detect sensitive data, it can't redact it. So, we can use option B With S3 Object Lambda and a prebuilt AWS Lambda function powered by Amazon Comprehend, you can protect PII data retrieved from S3 before returning it to an application.
7a1d491 👍 1 Selected: A
Amazon Macie is a fully managed data security and privacy service that uses machine learning to automatically discover, classify, and protect sensitive data, including personally identifiable information (PII). By running a sensitive data discovery job in Macie, the company can automatically identify PII in the S3 bucket and provide actionable insights to help secure it. The operational overhead is minimized because Macie handles the discovery and classification of PII automatically.
michele_scar 👍 2 Selected: B
https://docs.aws.amazon.com/AmazonS3/latest/userguide/tutorial-s3-object-lambda-redact-pii.html
kupo777 👍 3
Correct Answer: A Amazon Macie is designed specifically for discovering and protecting sensitive data within AWS environments. It automates the process of identifying PII in your S3 buckets, allowing you to create jobs that can regularly scan for and manage sensitive information. This approach minimizes manual effort and integrates well into existing workflows, providing ongoing protection without requiring additional infrastructure or complex setups.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

S3 Object Lambda is the correct choice because it allows you to add custom code to process data as it is retrieved from S3 without changing your application architecture. By attaching a Lambda function (powered by services like Amazon Comprehend or simple logic), you can detect and redact PII on-the-fly before the data reaches the consumer team. This approach ensures that the raw data remains intact in the source but never exposes PII to downstream analysts, fulfilling the security requirement with zero additional storage or pipeline management overhead.

Why the Other Options Are Wrong

Amazon Macie (Option A) is designed for discovery and classification of sensitive data; it alerts you to the presence of PII but does not have the capability to automatically delete or redact the data from the object itself. Option C involves Amazon Data Firehose, which requires setting up a streaming pipeline, increasing operational overhead significantly compared to direct access transformation. Option D suggests using AWS Glue DataBrew to create a separate copy of non-PII data, which doubles storage costs and introduces manual ETL maintenance, violating the 'least operational overhead' constraint.

Community Comment Notes

The community consensus strongly supports Option B, with many users noting that Macie is limited to detection. As user paali noted, "Macie will only detect sensitive data, it can't redact it," highlighting the critical distinction between finding data and protecting it. User michele_scar provided a direct link to the AWS tutorial for S3 Object Lambda PII redaction, confirming this is an official AWS recommended pattern for this specific scenario.

Official Reference

Exam Strategy

When asked about 'least operational overhead' for data protection, look for serverless, integrated solutions that modify data at the point of use rather than moving data to new pipelines or buckets. Always distinguish between 'discovery' services (like Macie) and 'transformation/protection' services (like Object Lambda).

Frequently Asked Questions

Why isn't Amazon Macie the right answer for removing PII?

Macie discovers and classifies sensitive data but does not have the functionality to automatically redact or remove that data from the S3 objects.

What is the operational overhead difference between S3 Object Lambda and DataBrew?

Object Lambda processes data on-demand with no infrastructure to manage, while DataBrew requires configuring jobs, scheduling, and managing output storage.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide