DP-100 — Frequently Asked Questions
Community-vetted answers to 20 common questions about this exam.
The correct authorization method is typically Account Key or SAS Token. While Managed Identity is the most secure and recommended method for production, exam questions often focus on the specific mechanics of using Account Keys for Azure Files, as it's a common point of confusion compared to Blob storage which more readily supports Managed Identity in all scenarios.
MLClient.from_config() relies on the DefaultAzureCredential chain. It attempts to authenticate using a sequence of methods, typically starting with Environment Variables, then Managed Identity (if running on Azure), and finally falling back to an interactive browser login or Azure CLI credentials if available locally.
Parameters are passed using the ${{inputs.parameter_name}} or ${{jobs.job_name.outputs.output_name}} syntax within the command string of the job definition. These variables are substituted at runtime. You must also define the inputs in the job's inputs dictionary and ensure the training script accepts them via command-line arguments (e.g., using argparse).
You must deploy the model to an endpoint configured with a Virtual Network (VNet) and Private Endpoint. This allows the scoring traffic to stay within the Azure backbone network. You will need to configure the workspace and the compute cluster for 'VNet injection' and ensure the necessary service tags (like AzureResourceManager, AzureActiveDirectory) are allowed in the Network Security Group (NSG) rules, as 'no egress' usually implies blocking public internet but allowing trusted Microsoft services.
In SDK v2, you configure this within the environment or compute definition of the job. For serverless (PaaS), you typically specify the instance_type (e.g., 'Standard_DS3_v2') and instance_count directly in the job configuration (e.g., job.compute = 'azureml:serverless' or by defining a specific cluster). If using a CommandJob, these are properties of the job object itself.
The correct order is: 1. Create the Batch Endpoint (the logical routing entity). 2. Create a Deployment (the compute and model configuration) and associate it with the endpoint. 3. Set the default deployment (optional, if you want one to handle traffic by default). 4. Invoke the endpoint with input data.
Random Sampling is a sampling method that selects hyperparameter values randomly from the defined search space. Early Termination (like Bandit or Median Stopping) is a policy that cancels low-performing runs before they complete to save compute resources. You can use Random Sampling to pick values and Early Termination to kill bad runs simultaneously.
You generally don't configure the kernel in the terminal; the terminal is a bash shell. However, to ensure a specific kernel is available for the Jupyter notebook interface on that Compute Instance, you install the ipykernel package in the specific Conda or Virtual environment you created. The kernel name in Jupyter will match the environment name.
You generally cannot 'update' the identity of an existing attached compute directly in a way that changes its underlying Azure resource ID. If you need to change the identity (e.g., from System-Assigned to User-Assigned), you typically must detach the Synapse pool from the Azure ML workspace and reattach it, specifying the new identity configuration during the attachment process.
You should use the MLTable asset. Define the path to the CSV and use the transformations (read_as_delimited) in the MLTable file or definition. Then, in your script, load it using mltable.load() which returns a Pandas DataFrame directly. This is preferred over using the older Dataset SDK methods.
The fewest properties required are the name (or the variable name in the script), the size (VM size, e.g., 'Standard_DS3_v2'), and the tier (usually 'Dedicated' or 'LowPriority'). The min_instances defaults to 0 and max_instances must be set (minimum 1).
You need the Subscription ID, Resource Group Name, and Workspace Name. You pass these along with a credential object (like DefaultAzureCredential) to the MLClient constructor: MLClient(credential, subscription_id, resource_group_name, workspace_name).
The client requires a valid Credential object (authentication), the Subscription ID, the Resource Group, and the Workspace Name. Without these four specific elements, the client cannot establish a connection to the control plane of the workspace.
You use the list() method on the client (e.g., client.components or client.jobs) and apply an OData filter string. For example: client.components.list(name='my_component') or using a filter query string like "name eq 'my_component'" depending on the specific SDK method version.
You typically need the azureml-ai-interpretability or azureml-ai-rai-insights (depending on the specific version/preview) along with the standard mlflow and azureml-mlflow libraries. The key is ensuring the model is logged with specific signatures that the RAI dashboard expects.
The three primary methods are Random, Grid, and Bayesian. Random picks random combinations, Grid tries every combination (expensive), and Bayesian uses previous run results to pick the next best set of parameters (intelligent).
You can log a dictionary as a set of parameters using mlflow.log_params(my_dict) or as metrics using mlflow.log_metrics(my_dict). If the dictionary contains complex objects or files, you should save it as a JSON or pickle file and log it as an artifact using mlflow.log_artifact('filename.json').
The common modes are mount (streams data to the compute, good for large data), download (copies data to local storage, good for small data), and upload (for outputs, copies data from compute to datastore). In SDK v2, these are defined in the Input and Output objects (e.g., Input(type=AssetTypes.URI_FOLDER, path=..., mode=InputOutputModes.MOUNT)).
1. Load the model using mlflow.<flavor>.load_model. 2. Create an OnlineEndpoint object. 3. Create an OnlineDeployment object, specifying the model (using the azureml:// URI of the registered model), instance_type, and code_configuration (pointing to the scoring script and environment). 4. Use client.begin_create_or_update(endpoint) and client.begin_create_or_update(deployment).
Authentication is handled via Key-based authentication (primary/secondary keys) or Azure Active Directory (Token-based) authentication. Cost monitoring is handled through Azure Cost Management by tagging the endpoint resource or viewing the associated Compute costs in the billing section of the Azure Portal, as the endpoint incurs charges based on the provisioned instance hours.
← Back to Microsoft DP-100 Exam Questions & Knowledge Points Guide