How to Dynamically Redact S3 PII for Multiple Applications?

Ensure data encryption and masking. Automate data processing by using AWS services. Understand data privacy and governance.
Answer Correct answer: B — Use an S3 Object Lambda endpoint to dynamically redact PII on S3 GET requests based on each application's needs, without creating multiple dataset copies.

A company has multiple applications that use datasets that are stored in an Amazon S3 bucket. The company has an ecommerce application that generates a dataset that contains personally identifiable information (PII). The company has an internal analytics application that does not require access to the PII. To comply with regulations, the company must not share PII unnecessarily. A data engineer needs to implement a solution that with redact PII dynamically, based on the needs of each application that accesses the dataset. Which solution will meet the requirements with the LEAST operational overhead?

  1. Create an S3 bucket policy to limit the access each application has. Create multiple copies of the dataset. Give each dataset copy the appropriate level of redaction for the needs of the application that accesses the copy.
  2. Create an S3 Object Lambda endpoint. Use the S3 Object Lambda endpoint to read data from the S3 bucket. Implement redaction logic within an S3 Object Lambda function to dynamically redact PII based on the needs of each application that accesses the data. Correct Answer
  3. Use AWS Glue to transform the data for each application. Create multiple copies of the dataset. Give each dataset copy the appropriate level of redaction for the needs of the application that accesses the copy.
  4. Create an API Gateway endpoint that has custom authorizers. Use the API Gateway endpoint to read data from the S3 bucket. Initiate a REST API call to dynamically redact PII based on the needs of each application that accesses the data.

Community Votes

B
100%

100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Tests dynamic PII redaction in S3 with minimal operational overhead; the trap is selecting bucket policies or Glue transformations that require multiple redacted copies.

Amazon S3 Object Lambda runs code on S3 GET requests so PII can be redacted per application without storing duplicate datasets. This page establishes why option B meets the dynamic redaction requirement with the least operational overhead on DEA-C01.

The most common wrong choice is A: using S3 bucket policies plus multiple copies of the dataset, because it appears to enforce least privilege but leaves PII in each copy and adds duplication and synchronization overhead.

Community Discussion (6 comments)

teo2157 👍 1 Selected: B
It's B based on AWS documentation https://docs.aws.amazon.com/AmazonS3/latest/userguide/transforming-objects.html
pypelyncar 👍 3 Selected: B
S3 Object Lambda automatically triggers the Lambda function only when there's a request to access data in the S3 bucket. This eliminates the need for pre-processing or creating multiple data copies with varying levels of redaction (Options A and C).
4c78df0 👍 2 Selected: B
B is correct
damaldon 👍 4
Ans. B You can use an Amazon S3 Object Lambda Access Point to control access to documents with personally identifiable information (PII). https://docs.aws.amazon.com/comprehend/latest/dg/using-access-points.html
atu1789 👍 1 Selected: B
S3 Object Lambda allows you to add custom processing, such as redaction of PII, to data retrieved from S3. This is done dynamically, meaning you don’t need to store multiple copies of the data. It’s a more efficient and operationally simpler approach compared to managing multiple dataset versions.
rralucard_ 👍 3 Selected: B
Amazon S3 Object Lambda allows you to add your own code to S3 GET requests to modify and process data as it is returned to an application. For example, you could use an S3 Object Lambda to dynamically redact personally identifiable information (PII) from data retrieved from S3. This would allow you to control access to sensitive information based on the needs of different applications, without having to create and manage multiple copies of your data.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Amazon S3 Object Lambda is designed to intercept S3 GET requests and run custom code before data is returned to the caller. Because the redaction logic lives in a Lambda function, the company can inspect the requesting application's identity and return only the fields that application is allowed to see, including fully redacted PII for the internal analytics application. The stored dataset remains a single object, so there is no need to create multiple redacted copies or maintain ETL jobs for each consumer. This directly satisfies the dynamic, per-application redaction requirement while keeping operational overhead low, which is why B is the DEA-C01 answer. The service is managed by AWS and scales with S3 access patterns, so the team does not operate additional infrastructure.

Why the Other Options Are Wrong

Option A uses an S3 bucket policy with multiple dataset copies; bucket policies can restrict who reads an object, but they do not redact PII inside the object, and storing one copy per redaction level creates duplicate data and ongoing sync work. Option C relies on AWS Glue to transform data for each application and likewise creates multiple copies, adding job orchestration and storage cost that the scenario's least-overhead requirement rules out. Option D introduces an API Gateway endpoint with custom authorizers, but API Gateway is not an S3 data-access redaction service; it would require custom REST API logic and still needs a compute layer to read, redact, and return objects. A and C also risk stale or inconsistently redacted copies, while D adds latency and an extra public endpoint without simplifying PII handling.

Community Comment Notes

Several learners converge on S3 Object Lambda for this scenario. damaldon cites the Amazon Comprehend documentation to show that an S3 Object Lambda Access Point can "control access to documents with personally identifiable information (PII)". pypelyncar notes that S3 Object Lambda triggers the Lambda function only when data is requested, and that "This eliminates the need for pre-processing or creating multiple data copies". rralucard_ explains that S3 Object Lambda lets you "add your own code to S3 GET requests to modify and process data" as it is returned to an application. teo2157 points to the AWS documentation on transforming objects, and atu1789 adds that dynamic redaction avoids managing multiple dataset versions.

Official Reference

Exam Strategy

When DEA-C01 asks for the least operational overhead, prefer a managed, request-time processing service over solutions that duplicate data or chain multiple services. S3 Object Lambda is the AWS-native answer when redaction or transformation must vary by caller without changing stored objects.

Frequently Asked Questions

Why not use S3 bucket policies with multiple redacted copies?

Bucket policies control access but do not redact PII inside objects, and multiple copies add storage, sync, and governance overhead that the least-overhead requirement rules out.

Does S3 Object Lambda redact PII per application?

Yes. It runs Lambda code on S3 GET requests, so you can identify the caller and return redacted data dynamically without modifying the stored dataset.

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide