Masking Credit Card Numbers in Daily CSV Files with AWS Glue Sensitive Data Detection

Ensure data integrity and prepare data for modeling.
Answer Correct answer: B — AWS Glue's Sensitive Data Detection finds and masks credit card numbers inside the existing Spark job, so no custom regex is needed and Macie cannot mask data.

A company receives daily .csv files about customer interactions with its ML model. The company stores the files in Amazon S3 and uses the files to retrain the model. An ML engineer needs to implement a solution to mask credit card numbers in the files before the model is retrained. Which solution will meet this requirement with the LEAST development effort?

  1. Create a discovery job in Amazon Macie. Configure the job to find and mask sensitive data.
  2. Create Apache Spark code to run on an AWS Glue job. Use the Sensitive Data Detection functionality in AWS Glue to find and mask sensitive data. Correct Answer
  3. Create Apache Spark code to run on an AWS Glue job. Program the code to perform a regex operation to find and mask sensitive data.
  4. Create Apache Spark code to run on an Amazon EC2 instance. Program the code to perform an operation to find and mask sensitive data.

Community Votes

B
67%
A
33%

67% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

AWS Glue includes a Sensitive Data Detection capability that finds and masks sensitive information such as credit card numbers inside the Spark job, so the company reuses a managed detection function instead of writing regex logic or running Macie, which cannot modify data.

A company receives daily CSV files of customer interactions in S3 and uses them to retrain an ML model, and credit card numbers must be masked before retraining with the least development effort. The masking must be applied to files sitting in S3 as part of a repeatable pipeline.

Believing an Amazon Macie discovery job can mask sensitive data. Macie is a discovery and classification service that reports where sensitive data lives; it has no action that rewrites or masks the values, so the masking requirement cannot be met with Macie.

Community Discussion (3 comments)

eesa 👍 1 Selected: B
La empresa necesita una solución automatizada para enmascarar números de tarjetas de crédito en archivos CSV almacenados en Amazon S3, antes de que los datos sean utilizados para reentrenar un modelo de Machine Learning. La opción B es la mejor solución porque: ✅ AWS Glue tiene integración nativa con S3, lo que facilita el procesamiento de los archivos sin desarrollo adicional. ✅ Glue Studio proporciona detección de datos sensibles (Sensitive Data Detection), lo que reduce el esfuerzo de desarrollo comparado con programar regex manualmente. ✅ Es una solución completamente administrada, eliminando la necesidad de mantener infraestructura en EC2. ✅ Spark en AWS Glue permite manejar grandes volúmenes de datos, ideal para procesamiento de múltiples archivos diarios.
michele_scar 👍 1 Selected: B
Macie can't take any actions, It can be only a discovery service.
ygn4ei 👍 1 Selected: A
correct

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

The requirement is to mask credit card numbers in the daily CSV files before retraining, with the least development effort. AWS Glue provides a Sensitive Data Detection functionality that is available inside the Spark job, so the company gets an already-built detector and masking routine rather than writing its own logic. Because the files already live in S3 and Glue reads them natively, the same job both processes and masks the data with minimal code. The vote was 67 for B, and eesa highlighted that Glue has native S3 integration and that Glue Studio provides the development experience, while the built-in sensitive data detection removes the need for custom masking code.

Why the Other Options Are Wrong

Creating a Macie discovery job to find and mask sensitive data (A) is not possible as described, because Macie is a discovery and classification service that identifies sensitive data but never modifies it. Writing Apache Spark code on an AWS Glue job that performs a regex to find and mask the data (C) does work technically, but it is strictly more development effort than using the Sensitive Data Detection functionality that Glue already provides for the same purpose, so it fails the least-development-effort requirement. Running the same custom Spark code on an Amazon EC2 instance (D) adds cluster management on top of the same hand-written regex, which is the highest-effort option of all.

Community Comment Notes

The community voted 67 for B and 33 for A. The decisive point came from michele_scar, who noted that Macie cannot take any actions and is only a discovery service, which directly invalidates the option A position. eesa's Spanish-language explanation reinforced option B by citing Glue's native S3 integration and the development experience in Glue Studio, and no commenter defended option A on the grounds that Macie could perform the masking itself.

Official Reference

Related Analysis

Practice All MLA-C01 Questions

Access 115 questions with complete answers and detailed explanations.

View Full MLA-C01 Practice Test →

← Back to MLA-C01 Study Guide