Bird
Raised Fist0
SciPydata~10 mins

Goodness of fit evaluation in SciPy - Step-by-Step Execution

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Concept Flow - Goodness of fit evaluation
Collect observed data
Define expected distribution
Calculate test statistic
Compare statistic to distribution
Get p-value
Decide if fit is good or not
We start with observed data and an expected distribution, calculate a test statistic, then find a p-value to decide if the data fits well.
Execution Sample
SciPy
from scipy.stats import chisquare
observed = [16, 18, 16, 14, 12, 12]
expected = [15, 15, 15, 15, 15, 15]
stat, p = chisquare(f_obs=observed, f_exp=expected)
print(stat, p)
This code runs a chi-square goodness of fit test comparing observed counts to expected counts.
Execution Table
StepActionCalculationResult
1Calculate differences (observed - expected)[16-15, 18-15, 16-15, 14-15, 12-15, 12-15][1, 3, 1, -1, -3, -3]
2Square differences[1^2, 3^2, 1^2, (-1)^2, (-3)^2, (-3)^2][1, 9, 1, 1, 9, 9]
3Divide squared differences by expected[1/15, 9/15, 1/15, 1/15, 9/15, 9/15][0.0667, 0.6, 0.0667, 0.0667, 0.6, 0.6]
4Sum all values0.0667+0.6+0.0667+0.0667+0.6+0.62.0
5Calculate p-value from chi-square distribution with df=5p = 1 - CDF(2.0, df=5)p = 0.849
6Decisionp > 0.05 means fit is goodFail to reject null hypothesis
💡 Test ends after p-value calculation and decision step
Variable Tracker
VariableStartAfter Step 1After Step 2After Step 3After Step 4After Step 5Final
observed[16,18,16,14,12,12][16,18,16,14,12,12][16,18,16,14,12,12][16,18,16,14,12,12][16,18,16,14,12,12][16,18,16,14,12,12][16,18,16,14,12,12]
expected[15,15,15,15,15,15][15,15,15,15,15,15][15,15,15,15,15,15][15,15,15,15,15,15][15,15,15,15,15,15][15,15,15,15,15,15][15,15,15,15,15,15]
diffN/A[1,3,1,-1,-3,-3][1,3,1,-1,-3,-3][1,3,1,-1,-3,-3][1,3,1,-1,-3,-3][1,3,1,-1,-3,-3][1,3,1,-1,-3,-3]
squared_diffN/AN/A[1,9,1,1,9,9][1,9,1,1,9,9][1,9,1,1,9,9][1,9,1,1,9,9][1,9,1,1,9,9]
chi_componentsN/AN/AN/A[0.0667,0.6,0.0667,0.0667,0.6,0.6][0.0667,0.6,0.0667,0.0667,0.6,0.6][0.0667,0.6,0.0667,0.0667,0.6,0.6][0.0667,0.6,0.0667,0.0667,0.6,0.6]
statisticN/AN/AN/AN/A2.02.02.0
p_valueN/AN/AN/AN/AN/A0.8490.849
Key Moments - 3 Insights
Why do we square the differences between observed and expected counts?
Squaring makes all differences positive and emphasizes larger differences, as shown in step 2 of the execution_table.
What does a high p-value mean in this test?
A high p-value (like 0.849 in step 5) means the observed data fits the expected distribution well, so we do not reject the fit.
Why do we divide squared differences by expected counts?
Dividing by expected counts normalizes the differences, so categories with larger expected counts don't dominate, as in step 3.
Visual Quiz - 3 Questions
Test your understanding
Look at the execution_table at step 4, what is the sum of the chi-square components?
A2.0
B1.5
C3.0
D0.85
💡 Hint
Check the 'Sum all values' calculation in step 4 of the execution_table.
According to variable_tracker, what is the value of p_value after step 5?
A0.15
B0.849
C0.05
D2.0
💡 Hint
Look at the 'p_value' row under 'After Step 5' in variable_tracker.
If the observed counts were all equal to expected counts, what would the chi-square statistic be?
A6
B1
C0
D15
💡 Hint
When observed equals expected, differences are zero, so sum of squared differences is zero (see step 1 and 2).
Concept Snapshot
Goodness of fit test compares observed data to expected distribution.
Calculate chi-square statistic: sum((observed - expected)^2 / expected).
Find p-value from chi-square distribution with degrees of freedom = categories - 1.
High p-value means data fits well; low p-value means poor fit.
Use scipy.stats.chisquare for easy calculation.
Full Transcript
Goodness of fit evaluation checks if observed data matches an expected pattern. We start by collecting observed counts and defining expected counts. Then, we calculate the differences between observed and expected, square them, and divide by expected counts. Summing these gives the chi-square statistic. Using the chi-square distribution with the right degrees of freedom, we find the p-value. A high p-value means the observed data fits the expected distribution well. This process is shown step-by-step in the execution table and variable tracker. Key points include squaring differences to avoid negatives, normalizing by expected counts, and interpreting the p-value correctly.

Practice

(1/5)
1. What does the chi-square goodness of fit test in scipy.stats.chisquare primarily evaluate?
easy
A. The mean difference between two samples
B. How well observed data matches expected frequencies
C. The correlation between two variables
D. The variance within a single dataset

Solution

  1. Step 1: Understand the purpose of chi-square test

    The chi-square goodness of fit test compares observed data frequencies to expected frequencies to check if they match.
  2. Step 2: Identify what scipy.stats.chisquare does

    This function calculates the chi-square statistic and p-value to evaluate the fit between observed and expected counts.
  3. Final Answer:

    How well observed data matches expected frequencies -> Option B
  4. Quick Check:

    Goodness of fit = observed vs expected match [OK]
Hint: Chi-square tests observed vs expected frequencies [OK]
Common Mistakes:
  • Confusing goodness of fit with correlation
  • Thinking it measures mean differences
  • Mixing variance analysis with goodness of fit
2. Which of the following is the correct way to import the chi-square goodness of fit test function from scipy?
easy
A. import scipy.chisquare
B. from scipy import chisquare
C. from scipy.stats import chisquare
D. import scipy.stats.chisquare as cs

Solution

  1. Step 1: Recall the module structure of scipy

    The chi-square test function is inside the stats submodule of scipy.
  2. Step 2: Identify correct import syntax

    The correct import is from scipy.stats import chisquare to directly access the function.
  3. Final Answer:

    from scipy.stats import chisquare -> Option C
  4. Quick Check:

    Correct import = from scipy.stats import chisquare [OK]
Hint: Import from scipy.stats for statistical tests [OK]
Common Mistakes:
  • Trying to import chisquare directly from scipy
  • Using incorrect module paths
  • Using alias without import
3. What will be the output of the following code?
from scipy.stats import chisquare
observed = [20, 30, 50]
expected = [25, 25, 50]
result = chisquare(f_obs=observed, f_exp=expected)
print(round(result.statistic, 2), round(result.pvalue, 3))
medium
A. 2.0 0.368
B. 3.0 0.223
C. 0.5 0.778
D. 1.0 0.607

Solution

  1. Step 1: Calculate chi-square statistic manually

    Chi-square = sum((observed - expected)^2 / expected) = ((20-25)^2/25) + ((30-25)^2/25) + ((50-50)^2/50) = (25/25)+(25/25)+0 = 1+1+0 = 2.0
  2. Step 2: Interpret p-value from scipy output

    Using scipy.stats.chisquare with these values gives a p-value around 0.368, indicating moderate fit.
  3. Final Answer:

    2.0 0.368 -> Option A
  4. Quick Check:

    Chi-square stat = 2.0, p-value ≈ 0.368 [OK]
Hint: Calculate chi-square stat then check p-value [OK]
Common Mistakes:
  • Forgetting to square differences
  • Dividing by wrong expected values
  • Mixing up statistic and p-value
4. Identify the error in this code snippet for performing a chi-square goodness of fit test:
from scipy.stats import chisquare
observed = [15, 25, 35]
expected = [20, 20]
result = chisquare(f_obs=observed, f_exp=expected)
print(result)
medium
A. Observed and expected arrays have different lengths
B. chisquare function is not imported correctly
C. Expected frequencies must be integers
D. Missing p-value extraction from result

Solution

  1. Step 1: Check input array lengths

    The observed array has 3 elements, but expected has only 2 elements, which is invalid for chi-square test.
  2. Step 2: Understand scipy requirement

    Both observed and expected arrays must be the same length to compare frequencies correctly.
  3. Final Answer:

    Observed and expected arrays have different lengths -> Option A
  4. Quick Check:

    Array length mismatch causes error [OK]
Hint: Observed and expected must be same length [OK]
Common Mistakes:
  • Ignoring length mismatch
  • Assuming expected must be integers
  • Thinking import or print is the error
5. You have observed counts of [40, 35, 25] for three categories. You expect them to be equally likely. Using scipy.stats.chisquare, what is the p-value indicating if the observed data fits the equal distribution? (Hint: expected counts are equal for all categories.)
hard
A. 0.789
B. 0.223
C. 0.456
D. 0.174

Solution

  1. Step 1: Calculate expected counts for equal distribution

    Total counts = 40+35+25 = 100. Expected counts = [100/3, 100/3, 100/3] ≈ [33.33, 33.33, 33.33].
  2. Step 2: Perform chi-square test with scipy

    Using chisquare(f_obs=[40,35,25], f_exp=[33.33,33.33,33.33]) gives a chi-square statistic ≈ 3.5 and p-value ≈ 0.174.
  3. Final Answer:

    0.174 -> Option D
  4. Quick Check:

    Unequal counts vs equal expected gives p-value ≈ 0.174 [OK]
Hint: Equal expected counts = total/number categories [OK]
Common Mistakes:
  • Using observed counts as expected
  • Not dividing total counts equally
  • Misinterpreting p-value significance