DP-203 — Frequently Asked Questions

Community-vetted answers to 15 common questions about this exam.

You can verify result set cache usage by querying the dynamic management view (DMV) sys.dm_pdw_exec_requests. By examining the result_set_cache_hit column for a specific request, you can determine if the query result was served from the cache (true) or if it required a full execution (false).

A Self-hosted Integration Runtime (SHIR) is required. The SHIR is a client agent that you install on a machine within your on-premises network. It facilitates secure data transfer between cloud data stores and on-premises data stores by acting as a bridge.

To track query start and end times, you should enable the ExecRequests diagnostic log setting. This log captures detailed information about every request executed against the SQL pool, including request_id, start_time, and end_time, which is essential for performance monitoring and auditing.

When using resource-specific mode for ADF diagnostic logs, the pipeline run data is stored in the ADFPipelineRun table in your Log Analytics workspace. This is distinct from the legacy mode, which sends all logs to a single AzureDiagnostics table.

No, an ARM template export only captures the factory configuration as it exists in the live, published service. It does not include any unpublished changes or pipelines that exist only in the development branch of your source control repository. To save an unpublished pipeline, you must use source control (like Git) or export the JSON directly from the ADF UI.

Yes, enabling Git integration allows you to save unpublished pipelines. When Git integration is configured, your pipeline definitions are stored as JSON files in the collaboration branch of your Git repository. You can save your work to this branch at any time without needing to publish the changes to the live Data Factory service.

The most effective way is to export the pipeline's JSON definition directly from the ADF user interface. This method captures the exact state of the pipeline in the development environment, including all activities and configurations, regardless of its published status. This JSON file can then be imported into another data factory or stored in a source control system.

A key consideration is that the start time for a scheduled trigger is always interpreted as UTC. You must convert your local time to Coordinated Universal Time (UTC) when setting the schedule to ensure the pipeline runs at the intended time, regardless of the time zone of the user who created the trigger.

Azure Synapse Link for SQL Server is implemented by enabling the change feed on the source SQL Server database and then using Azure Data Factory to ingest this change data into an Azure Data Lake Storage Gen2 account. Synapse Serverless SQL pools can then be used to query this data directly from the data lake.

Simple, deterministic SELECT queries are eligible for result set caching. Queries are not eligible if they involve non-deterministic functions (like GETDATE()), user-defined functions (UDFs), views that contain non-deterministic functions, or operations on external tables.

The 'Lookup' activity is typically used to capture the output of a stored procedure. While the 'Stored Procedure' activity can execute the procedure, the 'Lookup' activity is designed to retrieve and return the result set (e.g., a single value or a table) as an output that can be used by downstream activities in the pipeline.

The 'Publish' action deploys the resources from your collaboration branch (e.g., main or adf_publish) in the Git repository to the live Data Factory service. This action updates the factory with the latest approved changes, making them available for triggers and manual execution.

Lake Databases in Azure Synapse are organized into groups. This grouping method allows you to logically categorize and manage multiple lake databases within a single workspace, making it easier to handle large numbers of databases and control access permissions at a group level.

The data for a serverless SQL pool database is stored externally in Azure Data Lake Storage Gen2 (ADLS Gen2). The serverless SQL pool itself is a compute service that queries the data directly from the files in the data lake; it does not store the data within its own infrastructure.

In a live mode (without Git integration), you cannot save certain changes without publishing them. However, with Git integration enabled, you can save changes to your collaboration branch without publishing. The constraint primarily applies to live mode, where the 'Save' function is tied to the 'Publish' action to update the live service.

← Back to Microsoft DP-203 Tricky Questions & Common Mistakes