AWS Glue PII Detection and Obfuscation

Answer Correct answer: B — Use the Detect PII transform in AWS Glue Studio to identify and obfuscate PII, then ingest into S3.

A data engineer must use AWS services to ingest a dataset into an Amazon S3 data lake. The data engineer profiles the dataset and discovers that the dataset contains personally identifiable information (PII). The data engineer must implement a solution to profile the dataset and obfuscate the PII. Which solution will meet this requirement with the LEAST operational effort?

  1. Use an Amazon Kinesis Data Firehose delivery stream to process the dataset. Create an AWS Lambda transform function to identify the PII. Use an AWS SDK to obfuscate the PII. Set the S3 data lake as the target for the delivery stream.
  2. Use the Detect PII transform in AWS Glue Studio to identify the PII. Obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake. Correct Answer
  3. Use the Detect PII transform in AWS Glue Studio to identify the PII. Create a rule in AWS Glue Data Quality to obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake.
  4. Ingest the dataset into Amazon DynamoDB. Create an AWS Lambda function to identify and obfuscate the PII in the DynamoDB table and to transform the data. Use the same Lambda function to ingest the data into the S3 data lake.

Community Votes

B
69%
C
31%

69% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

The core concept tested is leveraging managed AWS Glue features versus custom code or orchestration layers, with the common trap being the misconception that Data Quality is required for obfuscation or that Lambda is needed for simple transformations.

This question addresses the least operational effort solution for profiling and obfuscating Personally Identifiable Information (PII) in an Amazon S3 data lake using AWS services. It establishes that AWS Glue Studio's built-in capabilities handle both detection and obfuscation natively.

Many learners incorrectly choose Option C because they assume AWS Glue Data Quality handles obfuscation; however, Data Quality is strictly for validation and monitoring, not data transformation or masking.

Community Discussion (23 comments)

milofficial 👍 12 Selected: B
How does Data Quality obfuscate PII? You can do this directly in Glue Studio: https://docs.aws.amazon.com/glue/latest/dg/detect-PII.html
Khooks 👍 5 Selected: B
Option C involves additional steps and complexity with creating rules in AWS Glue Data Quality, which adds more operational effort compared to directly using AWS Glue Studio's capabilities.
Kalyso 👍 1 Selected: B
Actually it is B. No need to create a rule in AWS Glue.
plutonash 👍 1 Selected: C
B. Use the Detect PII transform in AWS Glue Studio to identify the PII. Obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake. Detect PII transform only detects. Obfuscate the PII ok but how ? Answer C explain how
Udyan 👍 1 Selected: C
Why C is better than B: Obfuscation clarity: Option C explicitly mentions using a Glue Data Quality rule to obfuscate PII, while option B does not specify how obfuscation is implemented. Accuracy: Glue Data Quality provides a more structured way to handle obfuscation compared to relying solely on Glue Studio's PII detection. Thus, C is the most accurate and operationally efficient solution.
markill123 👍 1
The keyt
antun3ra 👍 2 Selected: B
B provides a streamlined, mostly visual approach using purpose-built tools for data processing and PII handling, making it the solution with the least operational effort.
portland 👍 1 Selected: C
https://aws.amazon.com/blogs/big-data/automated-data-governance-with-aws-glue-data-quality-sensitive-data-detection-and-aws-lake-formation/
qwertyuio 👍 2 Selected: B
https://docs.aws.amazon.com/glue/latest/dg/detect-PII.html
bakarys 👍 1 Selected: C
anwser is C
bigfoot1501 👍 3
I don't think we need to use much more services to fulfill these requirements. Just AWS Glue is enough, it can detect and obfuscate PII data already. Source: https://docs.aws.amazon.com/glue/latest/dg/detect-PII.html#choose-action-pii
VerRi 👍 3 Selected: C
We cannot directly handle PII with Glue Studio, and Glue Data Quality can be used to handle PII.
Just_Ninja 👍 1 Selected: A
A very easy was is to use the SDK to identify PII. https://docs.aws.amazon.com/code-library/latest/ug/comprehend_example_comprehend_DetectPiiEntities_section.html
kairosfc 👍 3 Selected: C
The transform Detect PII in AWS Glue Studio is specifically used to identify personally identifiable information (PII) within the data. It can detect and flag this information, but on its own, it does not perform the obfuscation or removal of these details. To effectively obfuscate or alter the identified PII, an additional transformation would be necessary. This could be accomplished in several ways, such as: Writing a custom script within the same AWS Glue job using Python or Scala to modify the PII data as needed. Using AWS Glue Data Quality, if available, to create rules that automatically obfuscate or modify the data identified as PII. AWS Glue Data Quality is a newer tool that helps improve data quality through rules and transformations, but whether it's needed will depend on the functionality's availability and the specificity of the obfuscation requirements
okechi 👍 2
Answer is option C. Period
arvehisa 👍 4 Selected: B
B is correct. C: glue data quality cannot obfuscate the PII D: need to write code but the question is the "LEAST operational effort"
certplan 👍 2
In python --- from awsglue.utils import getResolvedOptions from pyspark.context import SparkContext from awsglue.context import GlueContext from pyspark.sql import SparkSession # Initialize Spark session spark = SparkSession.builder \ .appName("Example Glue Job") \ .getOrCreate() # Initialize Glue context glueContext = GlueContext(SparkContext.getOrCreate()) # Retrieve Glue job arguments args = getResolvedOptions(sys.argv, ['JOB_NAME']) # Define your EMR step emr_step = [ { "Name": "My EMR Step", "ActionOnFailure": "CONTINUE", "HadoopJarStep": { "Jar": "s3://your-bucket/emr-scripts/your_script.jar", "Args": [ "arg1", "arg2" ] } } ] # Execute the EMR step response = glueContext.start_job_run(args['JOB_NAME'], job_run_args={'--extra-py-files': 'your_script.py'}) print(response)
certplan 👍 2
B. Utilizes AWS Glue Studio for PII detection, AWS Step Functions for orchestration, and S3 for storage. Glue Studio simplifies PII detection, and Step Functions can streamline the data pipeline orchestration, potentially reducing operational effort compared to option A. C. Similar to option B, but it additionally includes AWS Glue Data Quality for obfuscating PII. This might add a bit more complexity but can also streamline the process if Glue Data Quality offers convenient features for PII obfuscation.
jellybella 👍 4 Selected: B
AWS Glue Data Quality is a feature that automatically validates the quality of the data during a Glue job run, but it's not typically used for data obfuscation.
GiorgioGss 👍 1 Selected: C
https://dev.to/awscommunity-asean/validating-data-quality-with-aws-glue-databrew-4df4 https://docs.aws.amazon.com/glue/latest/dg/detect-PII.html
BartoszGolebiowski24 👍 1
I think this is A. We ingest data to s3 with a PPI transformation. We do not need to use glue, or step function here in that case.
rralucard_ 👍 2 Selected: C
Option C seems to be the best solution to meet the requirement with the least operational effort. It leverages AWS Glue Studio for PII detection, AWS Glue Data Quality for obfuscation, and AWS Step Functions for orchestration, minimizing the need for custom coding and manual processes.
TonyStark0122 👍 1
C. Use the Detect PII transform in AWS Glue Studio to identify the PII. Create a rule in AWS Glue Data Quality to obfuscate the PII. Use an AWS Step Functions state machine to orchestrate a data pipeline to ingest the data into the S3 data lake

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option B is the correct answer because AWS Glue Studio provides a native 'Detect PII' transform that can identify PII and automatically apply obfuscation actions (such as masking or tokenization) without writing custom code. By using this single transform within a Glue job, the engineer achieves both profiling and obfuscation with minimal configuration. The Step Functions orchestration is a standard pattern for pipeline management but does not add significant operational overhead compared to building custom Lambda functions.

Why the Other Options Are Wrong

Option A involves creating a Kinesis Firehose stream and a custom Lambda function, which requires coding, testing, and managing more infrastructure than a Glue job, thus violating the 'least operational effort' requirement. Option C incorrectly suggests using AWS Glue Data Quality for obfuscation; Data Quality is designed for validating data against rules and generating reports, not for modifying or masking data values. Option D proposes ingesting into DynamoDB first, which adds unnecessary complexity, storage costs, and development effort for a task that can be done directly in Glue before loading to S3.

Community Comment Notes

Community comments heavily support Option B, with many users noting that Glue Studio handles PII detection and obfuscation directly. As user milofficial noted, "How does Data Quality obfuscate PII? You can do this directly in Glue Studio," highlighting the key distinction between validation and transformation. User bigfoot1501 confirmed that "Just AWS Glue is enough, it can detect and obfuscate PII data already," referencing official documentation. Some users initially chose C due to confusion about Data Quality's role, but consensus clarified that it cannot perform obfuscation.

Official Reference

Exam Strategy

When asked for 'least operational effort', prioritize fully managed, purpose-built services over custom code (Lambda) or multi-step complex workflows. Familiarize yourself with the specific capabilities of AWS Glue Studio transforms, especially regarding PII handling, to avoid confusing them with Data Quality or other separate services.

Frequently Asked Questions

Can AWS Glue Data Quality obfuscate PII?

No. AWS Glue Data Quality is used for validating data quality and monitoring metrics, not for transforming or obfuscating data values.

Why is Lambda less efficient than Glue for this task?

Lambda requires custom code for detection and obfuscation logic, whereas Glue Studio offers pre-built, no-code transforms for PII handling, reducing development time.

More DEA-C01 FAQ →

Related Analysis

Practice All DEA-C01 Questions

Access 100 questions with complete answers and detailed explanations.

View Full DEA-C01 Practice Test →

← Back to DEA-C01 Study Guide