Evaluate PySpark DataFrame Statistics with df.show()
Note: This question is part of a series of questions that present the same scenario. Each question in the series contains a unique solution that might meet the stated goals. Some question sets might have more than one correct solution, while others might not have a correct solution. After you answer a question in this section, you will NOT be able to return to it. As a result, these questions will not appear in the review screen. You have a Fabric tenant that contains a new semantic model in OneLake. You use a Fabric notebook to read the data into a Spark DataFrame. You need to evaluate the data to calculate the min, max, mean, and standard deviation values for all the string and numeric columns. Solution: You use the following PySpark expression: df.show() Does this meet the goal?
Community Votes
100% of anonymous learners picked answer B. Votes are pick records left by other test-takers — they are not the verified answer.
Community Insight
The question tests PySpark DataFrame evaluation methods, and the common trap is confusing a data display function (show) with a statistical computation function (describe).
This page clarifies why the PySpark expression df.show() fails to calculate statistical summaries like min, max, mean, and standard deviation for a Spark DataFrame. It establishes that df.describe() or df.summary() are the correct methods for this goal.
Choosing Yes because df.show() successfully interacts with the DataFrame, failing to realize it only displays rows rather than computing statistical metrics.
Community Discussion (6 comments)
Comments & Corrections
No comments yet — spotted an error or have a note? Share it below.
Expert Analysis
Why the Answer Is Correct
The correct answer is B (No) becausedf.show simply prints the top rows of the DataFrame to the console in a tabular format. It does not perform any statistical computations such as calculating min, max, mean, or standard deviation. Therefore, using df.show completely fails to meet the stated goal of evaluating the data for these specific statistics.Why the Other Options Are Wrong
Option A (Yes) is incorrect because it assumesdf.show calculates statistical metrics. While show is frequently used to preview data during development, it lacks any underlying logic to aggregate and compute descriptive statistics across string and numeric columns.Community Comment Notes
Multiple commenters pointed out thatdf.show only displays data, as stilferx noted with "df.show - shows the data in the dataframe". The correct method to compute these statistics is df.describe, which SamuComqi and others highlighted by stating "The correct syntax is df.describe". Official Reference
Exam Strategy
Know the exact purpose of common PySpark DataFrame methods. If a question asks for statistics (min, max, mean, stddev), immediately look for describe() or summary(), not display methods like show() or print().
Related Analysis
Practice All DP-600 Questions
Access 115 questions with complete answers and detailed explanations.
View Full DP-600 Practice Test →