Masking Sensitive Customer Features with AWS Glue DataBrew While Preserving Ordered Tabular Structure
A company wants to develop an ML model by using tabular data from its customers. The data contains meaningful ordered features with sensitive information that should not be discarded. An ML engineer must ensure that the sensitive data is masked before another team starts to build the model. Which solution will meet these requirements?
Community Votes
77% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
AWS Glue DataBrew provides purpose-built data masking and pseudonymization transforms that redact sensitive values in place, preserving the column structure and row order of the tabular dataset so the downstream team can still model it.
A company holds tabular customer data with meaningful ordered features containing sensitive information that must not be discarded, and the sensitive values must be masked before a different team builds a model. The solution has to protect the sensitive data while keeping the feature structure that the modeling team depends on intact.
Picking Amazon Macie, which is a data discovery and classification service that finds sensitive data but does not mask or transform it. The requirement is to mask the values, not merely to discover them.
Community Discussion (5 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
AWS Glue DataBrew is a visual data preparation tool built for tabular data, and it includes masking and pseudonymization transforms that replace sensitive values while keeping the columns, types, and row order of the dataset intact. Because the ordered features must be preserved and only the sensitive values obscured, an in-place masking transform is exactly the right fit, and DataBrew's low-code interface also suits a handoff between teams. ZoroRoronoa noted that DataBrew is the most efficient and user-friendly option for masking sensitive information while retaining the structure and order of tabular data, and added the key clarification that Macie cannot mask data. GiorgioGss linked the DataBrew personal-information-protection documentation, and eesa emphasized that DataBrew keeps the order and structure of features intact.Why the Other Options Are Wrong
Amazon Macie (A) is a managed data discovery and classification service that identifies where sensitive data lives; it does not perform masking or pseudonymization, so it cannot satisfy a requirement to mask the values before another team uses them. An AWS Batch job changing sensitive data to random values (C) would destroy the meaningful ordered structure of the features and requires writing custom code, so it fails both the masking and the simplicity requirement. An Amazon EMR job replacing values with random values (D) has the same two problems as the Batch option, since it destroys feature semantics while adding the operational overhead of a managed cluster.Community Comment Notes
The vote was decisive at 77 for B and 23 for A, and the substance comments are unanimous. ZoroRoronoa, GiorgioGss, Saransundar, and eesa all selected Glue DataBrew and specifically framed it as masking or pseudonymizing sensitive data within tabular data while preserving feature order and structure, with GiorgioGss citing the official personal-information-protection page.Official Reference
Related Analysis
Practice All MLA-C01 Questions
Access 115 questions with complete answers and detailed explanations.
View Full MLA-C01 Practice Test →