Masking Sensitive Customer Features with AWS Glue DataBrew While Preserving Ordered Tabular Structure

Ensure data integrity and prepare data for modeling.
Answer Correct answer: B — AWS Glue DataBrew masks sensitive values in place with visual transforms, preserving the ordered tabular feature structure the other team needs for modeling.

A company wants to develop an ML model by using tabular data from its customers. The data contains meaningful ordered features with sensitive information that should not be discarded. An ML engineer must ensure that the sensitive data is masked before another team starts to build the model. Which solution will meet these requirements?

  1. Use Amazon Made to categorize the sensitive data.
  2. Prepare the data by using AWS Glue DataBrew. Correct Answer
  3. Run an AWS Batch job to change the sensitive data to random values.
  4. Run an Amazon EMR job to change the sensitive data to random values.

Community Votes

B
77%
A
23%

77% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

AWS Glue DataBrew provides purpose-built data masking and pseudonymization transforms that redact sensitive values in place, preserving the column structure and row order of the tabular dataset so the downstream team can still model it.

A company holds tabular customer data with meaningful ordered features containing sensitive information that must not be discarded, and the sensitive values must be masked before a different team builds a model. The solution has to protect the sensitive data while keeping the feature structure that the modeling team depends on intact.

Picking Amazon Macie, which is a data discovery and classification service that finds sensitive data but does not mask or transform it. The requirement is to mask the values, not merely to discover them.

Community Discussion (5 comments)

eesa 👍 1 Selected: B
Prepare the data by using AWS Glue DataBrew ✅ Purpose-built for data preparation — works with tabular data. ✅ Offers data masking, pseudonymization, and data transformations. ✅ No-code/low-code tool suitable for collaborative workflows between engineers and analysts. ✅ Keeps the order and structure of features intact. ✅ Best fit for masking sensitive info before ML modeling.
ZoroRoronoa 👍 4 Selected: B
AWS Glue DataBrew (Option B) is the most efficient and user-friendly solution for masking sensitive information while retaining the structure and order of tabular data, making it ideal for preparing data for ML model development. AWS Macie cannot mask data.
wizardmo 👍 3 Selected: A
question error made -> macie
Saransundar 👍 2 Selected: B
Customer tabular data → AWS Glue DataBrew → Mask sensitive data → Prepare for ML model building
GiorgioGss 👍 3 Selected: B
https://docs.aws.amazon.com/databrew/latest/dg/personal-information-protection.html

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

AWS Glue DataBrew is a visual data preparation tool built for tabular data, and it includes masking and pseudonymization transforms that replace sensitive values while keeping the columns, types, and row order of the dataset intact. Because the ordered features must be preserved and only the sensitive values obscured, an in-place masking transform is exactly the right fit, and DataBrew's low-code interface also suits a handoff between teams. ZoroRoronoa noted that DataBrew is the most efficient and user-friendly option for masking sensitive information while retaining the structure and order of tabular data, and added the key clarification that Macie cannot mask data. GiorgioGss linked the DataBrew personal-information-protection documentation, and eesa emphasized that DataBrew keeps the order and structure of features intact.

Why the Other Options Are Wrong

Amazon Macie (A) is a managed data discovery and classification service that identifies where sensitive data lives; it does not perform masking or pseudonymization, so it cannot satisfy a requirement to mask the values before another team uses them. An AWS Batch job changing sensitive data to random values (C) would destroy the meaningful ordered structure of the features and requires writing custom code, so it fails both the masking and the simplicity requirement. An Amazon EMR job replacing values with random values (D) has the same two problems as the Batch option, since it destroys feature semantics while adding the operational overhead of a managed cluster.

Community Comment Notes

The vote was decisive at 77 for B and 23 for A, and the substance comments are unanimous. ZoroRoronoa, GiorgioGss, Saransundar, and eesa all selected Glue DataBrew and specifically framed it as masking or pseudonymizing sensitive data within tabular data while preserving feature order and structure, with GiorgioGss citing the official personal-information-protection page.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide