PySpark DataFrame describe() for String and Numeric Stats

Answer Correct answer: B — df.describe() does not calculate mean and standard deviation for string columns, so it does not meet the goal for all columns.

Note: This question is part of a series of questions that present the same scenario. Each question in the series contains a unique solution that might meet the stated goals. Some question sets might have more than one correct solution, while others might not have a correct solution. After you answer a question in this section, you will NOT be able to return to it. As a result, these questions will not appear in the review screen. You have a Fabric tenant that contains a new semantic model in OneLake. You use a Fabric notebook to read the data into a Spark DataFrame. You need to evaluate the data to calculate the min, max, mean, and standard deviation values for all the string and numeric columns. Solution: You use the following PySpark expression: df.describe().show() Does this meet the goal?

  1. Yes
  2. No Correct Answer

Community Votes

A
69%
B
31%

69% of anonymous learners picked answer A. Votes are pick records left by other test-takers — they are not the verified answer.

Community Insight

Tests knowledge of df.describe() limitations: it excludes non-numeric data from statistical calculations like mean and stddev.

The PySpark df.describe() method computes statistics only for numeric columns. It does not calculate mean or standard deviation for string columns, making it insufficient for the stated goal.

Many assume describe() works for all column types because it returns count, min, and max for strings, ignoring that mean and stddev are undefined for text.

Community Discussion (9 comments)

Martin_Nbg 👍 14
I think A is correct https://learn.microsoft.com/en-us/dotnet/api/microsoft.spark.sql.dataframe.describe?view=spark-dotnet
Gunstsings 👍 5 Selected: A
DataFrame.Describe = Computes basic statistics for numeric and string columns, including count, mean, stddev, min, and max. If no columns are given, this function computes statistics for all numerical or string columns.
jackjack1 👍 1 Selected: A
Describe as other mentioned shows all fields asked for. if we were to use summary, we would have to specify what fields we want as shown here: https://learn.microsoft.com/en-us/dotnet/api/microsoft.spark.sql.dataframe.summary?view=spark-dotnet
nappi1 👍 1 Selected: A
df.describe().show() returns count, mean, stddev, min, max for all the columns (numeric and non numeric)
2fe10ed 👍 2 Selected: A
https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.DataFrame.describe.html
maia01 👍 4 Selected: B
describe: only numeric columns with limited summary summary: numeric and non-numeric, broader summary
patricck 👍 1
There's a difference in numeric and string values. It's applicable for the numeric values but not for the string values and the question mentions both data
zeeneuser 👍 4
A is correct, confirming after trying the command from Notebook. Displays count as well in addition to min, max, mean, and standard deviation.
moteruky 👍 3
It shows all stat for numeric values but shows only 3 stat for string(count,min and max, it doesnt account for mean and std)

Comments & Corrections

No comments yet — spotted an error or have a note? Share it below.

Log in to comment, report an error, or add a note about this question.

Submitted for moderation before publishing. Keep it helpful and respectful.

Expert Analysis

Why the Answer Is Correct

df.describe in PySpark is designed to compute basic statistics (count, mean, stddev, min, max) exclusively for numeric columns. For string columns, it typically only provides count, min, and max values. Since the requirement is to calculate mean and standard deviation for ALL string and numeric columns, this solution fails because those metrics cannot be meaningfully calculated for string data.

Why the Other Options Are Wrong

Option A claims the solution meets the goal, which is incorrect. While df.describe.show runs without error, it does not produce mean or standard deviation for string columns, thus failing the specific requirement for all columns.

Community Comment Notes

Several users like Gunstsings and Martin_Nbg incorrectly believe describe handles all stats for all columns. Others like maia01 and moteruky correctly identify that string columns lack mean/stddev support, confirming the answer is No. patricck also highlights the distinction between numeric and string applicability.

Official Reference

Exam Strategy

Always distinguish between descriptive statistics applicable to numeric vs. categorical data. If a question asks for mean/stddev on strings, the answer is almost always 'No' unless a custom UDF is specified.

Frequently Asked Questions

Why doesn't df.describe() work for string columns?

It only computes count, min, and max for strings. Mean and stddev are mathematically undefined for non-numeric data.

How to get stats for all columns in Spark?

Use summary() with explicit aggregation functions or filter columns by type before applying describe().

Related Analysis

Practice All DP-600 Questions

Access 115 questions with complete answers and detailed explanations.

View Full DP-600 Practice Test →

← Back to DP-600 Study Guide