How to Secure an ML Pipeline with Sensitive BigQuery Data?
Your organization is developing a sophisticated machine learning (ML) model to predict customer behavior for targeted marketing campaigns. The BigQuery dataset used for training includes sensitive personal information. You must design the security controls around the AI/ML pipeline. Data privacy must be maintained throughout the model’s lifecycle and you must ensure that personal data is not used in the training process. Additionally, you must restrict access to the dataset to an authorized subset of people only. What should you do?
Community Votes
100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
Tests the ability to apply data privacy controls specifically for ML training workloads, where candidates often confuse encryption or network proxies with actual data de-identification.
Securing machine learning pipelines with sensitive datasets requires proactive data sanitization and strict access governance. This page confirms that de-identifying data via Cloud DLP combined with IAM is the correct approach to prevent PII exposure during model training.
Option C is frequently chosen because customer-managed keys are commonly associated with data protection, but they only secure data at rest and do not remove sensitive information from the training dataset.
Community Discussion (3 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option A directly satisfies the requirement to ensure personal data is excluded from the training process by applying Cloud Data Loss Prevention (DLP) APIs to de-identify the BigQuery dataset beforehand. De-identification transforms or masks Personally Identifiable Information (PII), guaranteeing that sensitive attributes cannot be memorized or leaked by the ML model. Coupling this with strict IAM policies ensures that only authorized personnel can interact with the underlying data and pipeline resources, fulfilling both privacy and access control mandates.Why the Other Options Are Wrong
Option B relies on Identity-Aware Proxy, which secures application endpoints and network traffic but does not alter or protect the dataset contents during the training phase. Option C implements Customer-Managed Encryption Keys (CMEK), which provides robust at-rest encryption but leaves raw sensitive data fully accessible to users and services with BigQuery permissions, violating the no personal data in training rule. Option D deploys Confidential VMs to protect compute workloads in use, yet it fails to address the fundamental risk of feeding unredacted PII into the model training algorithm.Community Comment Notes
Community contributors consistently reinforce that de-identification is the only mechanism among the choices that actively removes sensitive information before it enters the training loop. As one learner noted, options focusing solely on encryption or network proxies "says nothing about data privacy" during the actual model development phase. The consensus correctly highlights that preserving privacy throughout the entire lifecycle requires proactive sanitization rather than passive storage or transmission safeguards.Official Reference
Exam Strategy
When securing AI/ML pipelines, always differentiate between protecting data in transit/rest and protecting data during processing. For training workloads involving PII, prioritize data sanitization techniques like masking, tokenization, or differential privacy over cryptographic controls alone.
Frequently Asked Questions
Why isn't CMEK sufficient for protecting training data?
CMEK encrypts data at rest but does not remove or mask PII, meaning sensitive information remains available to authorized users and can still be ingested into the ML training process.
How does Cloud DLP differ from manual data masking for AI?
DLP automatically discovers and classifies sensitive data before it enters the workflow, enabling programmatic de-identification tailored specifically for model training requirements without manual intervention.