How to Minimize Data Shuffling in PySpark Joins
You are analyzing customer purchases in a Fabric notebook by using PySpark. You have the following DataFrames: transactions: Contains five columns named transaction_id, customer_id, product_id, amount, and date and has 10 million rows, with each row representing a transaction. customers: Contains customer details in 1,000 rows and three columns named customer_id, name, and country. You need to join the DataFrames on the customer_id column. The solution must minimize data shuffling. You write the following code. from pyspark.sql import functions as F results = Which code should you run to populate the results DataFrame?
Community Votes
100% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests PySpark join optimization strategies, where the common trap is using a standard join which forces shuffling instead of broadcasting the smaller DataFrame.
This page explains how to minimize data shuffling when joining a large DataFrame with a small DataFrame in PySpark. It establishes that using the broadcast join optimization is the correct approach to avoid network shuffles.
Choosing a standard join (Option C) because it looks simpler, which fails to explicitly minimize shuffling and relies on Spark's potentially suboptimal auto-optimization.
Community Discussion (5 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
Option A usesF.broadcast(customers) to explicitly broadcast the smaller DataFrame (1,000 rows) to all worker nodes. This allows each node to perform the join locally with its partition of the larger DataFrame (10 million rows), completely avoiding the expensive network shuffle required by standard joins.Why the Other Options Are Wrong
Option C performs a standard join, which defaults to a SortMergeJoin that requires shuffling both DataFrames across the network to ensure matching keys are on the same node. Option B adds a.distinct operation, which introduces even more shuffling to resolve duplicates. Option D uses a crossJoin, which creates a massive Cartesian product (10 billion rows) before filtering, causing extreme performance degradation and shuffling.Community Comment Notes
Users correctly identified that broadcasting avoids shuffles by replicating the smaller table to all nodes, as Momoanwar noted: "avoids the need for network shuffles for each row of the larger table". sraakesh95 also pointed out that broadcasting "won't require any I/Os from other nodes, thereby, reducing the shuffling requirement".Official Reference
Exam Strategy
When joining a very large DataFrame with a small DataFrame in PySpark, always look for the broadcast join option to minimize shuffling. Remember that explicit broadcast hints override Spark's default join strategies and guarantee shuffle avoidance.
Related Analysis
Practice All DP-600 Questions
Access 115 questions with complete answers and detailed explanations.
View Full DP-600 Practice Test →