How to Secure an ML Pipeline with Sensitive BigQuery Data?

Data Privacy & Access Control
Answer Correct answer: A — De-identify sensitive BigQuery data using Cloud DLP APIs before model training and enforce strict IAM policies to restrict dataset access.

Your organization is developing a sophisticated machine learning (ML) model to predict customer behavior for targeted marketing campaigns. The BigQuery dataset used for training includes sensitive personal information. You must design the security controls around the AI/ML pipeline. Data privacy must be maintained throughout the model’s lifecycle and you must ensure that personal data is not used in the training process. Additionally, you must restrict access to the dataset to an authorized subset of people only. What should you do?

  1. De-identify sensitive data before model training by using Cloud Data Loss Prevention (DLP)APIs. and implement strict Identity and Access Management (IAM) policies to control access to BigQuery. Correct Answer
  2. Implement Identity-Aware Proxy to enforce context-aware access to BigQuery and models based on user identity and device.
  3. Implement at-rest encryption by using customer-managed encryption keys (CMEK) for the pipeline. Implement strict Identity and Access Management (IAM) policies to control access to BigQuery.
  4. Deploy the model on Confidential VMs for enhanced protection of data and code while in use. Implement strict Identity and Access Management (IAM) policies to control access to BigQuery.

Community Votes

A
100%

100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Tests the ability to apply data privacy controls specifically for ML training workloads, where candidates often confuse encryption or network proxies with actual data de-identification.

Securing machine learning pipelines with sensitive datasets requires proactive data sanitization and strict access governance. This page confirms that de-identifying data via Cloud DLP combined with IAM is the correct approach to prevent PII exposure during model training.

Option C is frequently chosen because customer-managed keys are commonly associated with data protection, but they only secure data at rest and do not remove sensitive information from the training dataset.

Community Discussion (3 comments)

532b5da 👍 1 Selected: A
Ans is A We want data privacy through out lifecycle. C is at rest D is in use B says nothing about data privacy
json4u 👍 1 Selected: A
It's A Well explained below.
abdelrahman89 👍 1
A - Data De-identification: De-identifying sensitive data using Cloud DLP APIs ensures that the data used for model training does not contain personally identifiable information (PII). This protects data privacy and reduces the risk of unauthorized access or misuse. IAM Policies: Implementing strict IAM policies controls access to BigQuery, ensuring that only authorized personnel can access and use the dataset. This further protects data privacy and reduces the risk of unauthorized access. Comprehensive Approach: This approach combines data de-identification and IAM controls to provide a robust and effective security solution for the AI/ML pipeline.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option A directly satisfies the requirement to ensure personal data is excluded from the training process by applying Cloud Data Loss Prevention (DLP) APIs to de-identify the BigQuery dataset beforehand. De-identification transforms or masks Personally Identifiable Information (PII), guaranteeing that sensitive attributes cannot be memorized or leaked by the ML model. Coupling this with strict IAM policies ensures that only authorized personnel can interact with the underlying data and pipeline resources, fulfilling both privacy and access control mandates.

Why the Other Options Are Wrong

Option B relies on Identity-Aware Proxy, which secures application endpoints and network traffic but does not alter or protect the dataset contents during the training phase. Option C implements Customer-Managed Encryption Keys (CMEK), which provides robust at-rest encryption but leaves raw sensitive data fully accessible to users and services with BigQuery permissions, violating the no personal data in training rule. Option D deploys Confidential VMs to protect compute workloads in use, yet it fails to address the fundamental risk of feeding unredacted PII into the model training algorithm.

Community Comment Notes

Community contributors consistently reinforce that de-identification is the only mechanism among the choices that actively removes sensitive information before it enters the training loop. As one learner noted, options focusing solely on encryption or network proxies "says nothing about data privacy" during the actual model development phase. The consensus correctly highlights that preserving privacy throughout the entire lifecycle requires proactive sanitization rather than passive storage or transmission safeguards.

Official Reference

Exam Strategy

When securing AI/ML pipelines, always differentiate between protecting data in transit/rest and protecting data during processing. For training workloads involving PII, prioritize data sanitization techniques like masking, tokenization, or differential privacy over cryptographic controls alone.

Frequently Asked Questions

Why isn't CMEK sufficient for protecting training data?

CMEK encrypts data at rest but does not remove or mask PII, meaning sensitive information remains available to authorized users and can still be ingested into the ML training process.

How does Cloud DLP differ from manual data masking for AI?

DLP automatically discovers and classifies sensitive data before it enters the workflow, enabling programmatic de-identification tailored specifically for model training requirements without manual intervention.

Related Analysis

← Back to PCSE Study Guide