How to Configure Spark Pool Access to External Data Lake Storage?

You have an Azure subscription that contains an Azure Data Lake Storage account named dl1 and an Azure Analytics Synapse workspace named workspace1. You need to query the data in dl1 by using an Apache Spark pool named Pool1 in workspace1. The solution must ensure that the data is accessible Pool1. Which two actions achieve the goal? Each correct answer presents a complete solution. NOTE: Each correct answer is worth one point.

  1. Implement Azure Synapse Link.
  2. Load the data to the primary storage account of workspace1. Source Reference Answer
  3. From workspace1, create a linked service for the dl1. Source Reference Answer
  4. From Microsoft Purview, register dl1 as a data source.

Community Votes

BC
100%

100% of anonymous learners picked answer BC. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

It tests how to grant Spark pool access to external ADLS Gen2 storage, commonly confused with governance tools like Microsoft Purview or unrelated features like Synapse Link.

This question tests configuring external data access for Azure Synapse Spark pools, with the community consensus confirming that linking the Data Lake Storage as a managed service or setting it as primary storage are required steps.

Option D (Microsoft Purview) is frequently chosen incorrectly because candidates assume registering a data source enables access, but Purview only handles metadata, lineage, and governance, not runtime connectivity.

Community Discussion (6 comments)

jongert 👍 13 Selected: BC
Would say Purview only registers the data source to track lineage, should not have anything to do with access. Synapse Link is not concerned with data lake. Seemingly, data lake storage has to be configured as primary storage as per: https://learn.microsoft.com/en-us/troubleshoot/azure/synapse-analytics/spark/spark-jobexec-storage-access#common-issues-and-solutions
Danweo 👍 1 Selected: BC
B and C, Purview doesnt give access
e56bb91 👍 1 Selected: BC
B works, C also works
mghf61 👍 1
A, C ?
be8a152 👍 1
C,D is the correct answer
dakku987 👍 1 Selected: BC
chatgpt(although not sure) C. From workspace1, create a linked service for dl1. D. From Microsoft Purview, register dl1 as a data source. Explanation: Create a linked service (C): You need to create a linked service in workspace1 that establishes a connection to dl1. This linked service allows your Spark pool (Pool1) to access the data in dl1. Register dl1 as a data source (D): Registering dl1 as a data source in Microsoft Purview helps in tracking lineage and metadata. While it is not directly related to enabling access from Spark, it is a good practice for governance and understanding the data landscape within your organization.

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

Option B and C correctly address how an Azure Synapse Spark pool accesses external storage. Configuring the Data Lake Storage account as the workspace’s primary storage automatically mounts it to the Spark cluster, granting immediate read/write permissions. Alternatively, creating a linked service establishes a secure, authenticated connection endpoint that Pool1 can reference to query dl1 without duplicating data. Both methods satisfy the requirement for runtime accessibility.

Why the Other Options Are Wrong

Option A is incorrect because Azure Synapse Link specifically integrates Azure Cosmos DB with Synapse for analytical processing, not ADLS Gen2. Option D is a frequent distractor; Microsoft Purview is strictly a data governance and cataloging tool that tracks lineage and metadata. Registering a datastore in Purview does not provision compute resources or configure authentication for Spark jobs to actually read the files.

Community Comment Notes

Candidates overwhelmingly selected BC, aligning with official Microsoft guidance. Comment [1] correctly points out that Purview handles lineage tracking rather than access control, and cites the official troubleshooting documentation for Spark storage access. Comment [2] reinforces that governance tools do not enable compute-to-storage connectivity. The consensus confirms that linking or mounting the storage account is the only viable path for Spark pool access.

Official Reference

Exam Strategy

Focus on distinguishing between data connectivity configurations and data governance tools when designing Azure Synapse solutions. Always verify whether a requirement involves runtime compute access versus metadata tracking before selecting options involving Purview or cataloging services.

Related Analysis

← Back to DP-203 Study Guide