The Bonferroni Betrayal: How I Fixed a Repeated Measures ANOVA in SPSS for a DNP Student

The Bonferroni Betrayal: How I Fixed a Repeated Measures ANOVA in SPSS for a DNP Student

Repeated measures ANOVA in SPSS is not the same as running three paired t-tests back to back, and that mix-up nearly cost a Doctor of Nursing Practice (DNP) student her project defense. I got the call five weeks before the defense date. The advisor had already signed off. The slides were half made.

This is a real case from my consulting work. Identifying details about the program and institution have been changed. This is also the clearest example I have of why a “correct-looking” SPSS output can still be wrong.

What Went Wrong Five Weeks Before the Defense

The student, working in a nurse anesthesia program, had delivered a difficult-airway simulation to 34 student registered nurse anesthetists (SRNAs). Confidence was measured on a 100-point scale at three points: before the simulation, right after, and at six-week follow-up.

The plan was simple. Confidence would rise after training and stay high at follow-up.

To test this, the student ran three separate paired samples t-tests in SPSS:

  • Baseline vs post-simulation: p = .001
  • Post-simulation vs follow-up: p = .18
  • Baseline vs follow-up: p = .04

The conclusion looked clean: confidence rose, held steady, and stayed elevated at six weeks. I asked one question before looking at anything else: did she run a repeated measures ANOVA in SPSS, or three separate t-tests?

Three separate t-tests, and that was mistake one, with three more waiting underneath it.

Problem One: Three T-Tests Are Not a Repeated Measures ANOVA in SPSS

When the same people are measured three times, the scores are not independent of each other. A person’s post-simulation confidence is tied to their own baseline. Running three separate paired t-tests ignores that link and creates a multiple comparisons problem.

The Familywise Error Rate Math

Every hypothesis test carries a 5% chance of a false positive. Run three tests on the same data and the combined error rate is not 5%. It works out to 1 − (0.95)³, which is roughly 14.3% — close to three times the risk the student thought she was taking.

One fix is the Bonferroni correction: divide alpha by the number of comparisons. With three tests, the new threshold is .05 / 3 = .0167. Under that line, the baseline-vs-follow-up result (p = .04) stops being significant. The claim that confidence stayed elevated at six weeks would not survive.

Bonferroni is a blunt tool though, protecting you from false positives while making it easier to miss a real effect. For a design with repeated measurements on the same people, the correct comparison against a paired t-test is a proper repeated measures ANOVA in SPSS: one omnibus test first, then post-hoc comparisons only if that omnibus test is significant.

What If Your Advisor Already Approved the Wrong Analysis?

This is not a reason to panic, and it is not something to hide. It happens more often than most programs admit, usually because the advisor was trained on between-subjects designs and the student’s project turned out to be within-subjects. A corrected analysis that supports the same overall story is a stronger showing in front of a committee, not a weaker one. In my experience, committees respond better to a student who caught and fixed the error than to a student who never noticed it existed.

Fixing this rarely means starting over. In this case, re-running the analysis and rewriting the results section took about a week, well inside a five-week runway, and it usually only takes a focused review of the affected tests, not a full redo of the whole project. Catching this before the defense, even late, is far cheaper than catching it during one.

How Do You Run a Repeated Measures ANOVA in SPSS?

Here is the syntax I had the student run instead of the three t-tests:

spss

GLM conf_base conf_post conf_follow
  /WSFACTOR=time 3 Polynomial
  /METHOD=SSTYPE(3)
  /POSTHOC=time (BONFERRONI)
  /PRINT=DESCRIPTIVE ETASQ OPOWER
  /CRITERIA=ALPHA(.05)
  /WSDESIGN=time.

This single command produces the multivariate tests, Mauchly’s test of sphericity, the within-subjects effects table, and Bonferroni-corrected pairwise comparisons together. The student had never seen Mauchly’s test before this; most people who learn SPSS from a quick tutorial never do, because most tutorials stop at the menu clicks.

Problem Two: What Happens When Sphericity Is Violated?

Sphericity is the repeated measures version of the equal-variance assumption you check in a between-subjects ANOVA. It assumes the variance of the differences between every pair of conditions is roughly equal. When it is not, the F-test gets too generous and hands you significant results that are not really there.

Sphericity, in plain terms: it is the assumption, in a within-subjects design, that the variances of the differences between all pairs of time points or conditions are equal. Mauchly’s test checks whether that assumption holds.

For Mauchly’s test of sphericity interpretation, look at one line: a chi-square value and a p-value. Here it read χ²(2) = 9.41, p = .009. Below .05 means sphericity is violated, which is exactly what happened.

Greenhouse-Geisser Correction SPSS: Which Row to Read

Once Mauchly’s test flags a violation, stop reading the “Sphericity Assumed” row of the within-subjects effects table. Move to a corrected row instead:

  1. Greenhouse-Geisser – the most conservative and most commonly reported correction
  2. Huynh-Feldt – slightly less conservative, useful for mild violations
  3. Lower-bound – the strictest option, rarely used in practice

I had the student read the Greenhouse-Geisser row. The epsilon value came out to .78, pulling the degrees of freedom down by 22%. That single adjustment widened the confidence intervals and shifted the p-values enough to matter.

The corrected omnibus result: F(1.56, 51.5) = 12.84, p < .001, partial η² = .28. In APA 7th edition style, this is the exact line a committee expects to see: F-value with corrected degrees of freedom, significance, and effect size, together. Confidence genuinely did change across the three time points, and the next question was where.

If sphericity is badly violated (epsilon well below .70) or your data are ordinal rather than interval, a correction is not always enough. The non-parametric alternative is the Friedman test, and it is worth running as a sensitivity check even when a correction technically clears the bar. Here, epsilon of .78 and clean normality meant Greenhouse-Geisser was sufficient, but you can run your data through a Friedman test calculator before deciding.

Problem Three: What the Post-Hoc Comparisons Actually Showed

With the omnibus test significant, we moved to Bonferroni-corrected pairwise comparisons:

  • Baseline vs post: mean difference = 18.3, p < .001, 95% CI [10.2, 26.4]
  • Post vs follow-up: mean difference = −4.1, p = .31, 95% CI [−11.8, 3.6]
  • Baseline vs follow-up: mean difference = 14.2, p = .002, 95% CI [5.4, 23.0]

The original story held up. Confidence rose sharply after training, stayed roughly stable to six weeks, and remained significantly above baseline at follow-up. What changed was the width of the confidence intervals and, more importantly, the effect size that got reported.

Partial Eta Squared vs Cohen’s d

The student had originally reported Cohen’s d for each of the three t-tests. That is a mistake in a repeated measures design, because Cohen’s d is not built for three related conditions. The correct measure here is partial eta squared (η²ₚ), which came out to .28 for the effect of time — large by Cohen’s own conventions (.14 and above is large).

Partial eta squared has a limitation though. It can look inflated in repeated measures designs because it does not account for variance sitting between subjects. I recommended reporting the more conservative generalized eta squared alongside it, available in SPSS through the EMMEANS subcommand, so the committee sees both numbers and understands why they differ. You can check your effect size against a calculator before you write it into your results section.

Problem Four: The Missing Data Was Handled Wrong

While going through the raw data file, I found something the three t-tests had quietly hidden. Two participants had missing scores at the six-week follow-up. SPSS had used listwise deletion by default, dropping both from every analysis, not just the follow-up comparison.

That took the sample from 34 down to 32. The two who dropped out were not a random subset either — their baseline confidence scores were the two lowest in the group (41 and 38). Removing them pulled the baseline mean upward and made the improvement from baseline to follow-up look smaller than it really was.

How to Handle Missing Data in Repeated Measures SPSS

When data are missing at random, multiple imputation is the standard fix. When you want to use every available data point without imputing anything, a linear mixed model does the job instead. Here, a mixed model made more sense because dropout was tied to a real variable (low baseline confidence), not pure chance.

spss

MIXED confidence BY time
  /FIXED=time | SSTYPE(3)
  /RANDOM=INTERCEPT | SUBJECT(id) COVTYPE(VC)
  /REPEATED=time | SUBJECT(id) COVTYPE(AR1)
  /PRINT=SOLUTION TESTCOV
  /EMMEANS=TABLES(time) COMPARE ADJ(BONFERRONI).

The mixed model kept all 34 participants and gave slightly wider confidence intervals, but the pattern did not change. I recommended the mixed model as the primary analysis, with the repeated measures ANOVA reported as a sensitivity check. The committee accepted both.

I have seen the same pattern in a public health dissertation where three participants dropped out of a nutrition intervention at final follow-up, all three with the highest baseline BMI in the sample. Same lesson: check who left before deciding how to handle who is missing.

Paired Samples T-Test vs Repeated Measures ANOVA: The Real Difference

This is the comparison every student asks me about. A paired samples t-test is built for exactly two related measurements on the same people. The moment you have three or more time points or conditions on the same subjects, you need a repeated measures ANOVA in SPSS instead, not three or four separate t-tests strung together.

The difference is not stylistic. Multiple t-tests on related data inflate your Type I error, ignore sphericity entirely, and give you no single omnibus test to defend to a committee.

Where Most SPSS Tutorials Fall Short

Most guides on this topic walk you through Analyze, General Linear Model, Repeated Measures, and stop there — the standard menu path used across nearly every SPSS tutorial site. That path gets you an output. It does not tell you what to do when Mauchly’s test comes back significant, or flag that your default deletion setting just dropped your two lowest scorers.

A step-by-step click guide is not wrong; it is just incomplete. The American Psychological Association’s Publication Manual (7th edition) sets the reporting standard this output needs to match, and Andy Field’s Discovering Statistics Using IBM SPSS Statistics is one reference I point students to for a fuller explanation of why Greenhouse-Geisser adjusts degrees of freedom the way it does. Neither gets mentioned in the average click-through tutorial.

What Changed in the Final Project

  • Three separate t-tests replaced with a single repeated measures ANOVA and Bonferroni-corrected post-hoc tests
  • Mauchly’s test reported, with Greenhouse-Geisser correction applied because sphericity was violated
  • Effect size reported as partial eta squared and generalized eta squared, not Cohen’s d
  • Listwise deletion replaced with a linear mixed model using all 34 participants
  • Results section rewritten to match the corrected analysis

The student’s own reaction after seeing the corrected output stuck with me. She had thought repeated measures just meant running the same test three times, and nobody had told her the three scores were not independent of each other.

If You Are Looking for Statistics Help for DNP Project Work

If you are staring down a similar defense date with an analysis you are not fully sure of, you are not alone. Statistics help for DNP project work usually comes down to the same handful of checks: the right test for the design, the right correction when assumptions fail, and honest handling of missing data. Before touching SPSS again, it is worth checking your data for normality and confirming you are in repeated measures territory to begin with. If you are not even sure outside help is the right call, here’s how to tell if you need a dissertation expert.

If you would rather have a second pair of eyes on your actual output before your defense, that is exactly the kind of work I do through one-on-one SPSS data analysis support.

FAQ

u003cstrongu003eWhat is Mauchly’s test of sphericity and why does it matter?u003c/strongu003e

Mauchly’s test checks whether the variances of the differences between your repeated measures are equal. If it comes back significant (p u0026lt; .05), sphericity is violated and you need a corrected F-test, usually Greenhouse-Geisser, instead of the uncorrected row.u003cbru003e

u003cstrongu003eWhat’s the difference between a paired samples t-test and a repeated measures ANOVA?u003c/strongu003e

A paired t-test compares exactly two related measurements. A repeated measures ANOVA compares three or more related measurements on the same subjects and controls the overall error rate in a way separate t-tests cannot.u003cbru003e

u003cstrongu003eWhat do you do if sphericity is violated in SPSS?u003c/strongu003e

Read the Greenhouse-Geisser corrected row in the within-subjects effects table instead of the sphericity-assumed row. Huynh-Feldt is a milder alternative for small violations, and the Friedman test is the non-parametric fallback for severe ones.u003cbru003e

u003cstrongu003eHow do you handle missing data in a repeated measures design?u003c/strongu003e

Use multiple imputation when data are missing at random, or a linear mixed model (SPSS MIXED procedure) when you want to keep every participant without imputing values. Always check whether dropout is related to another variable, such as baseline score, before choosing.u003cbru003e

u003cstrongu003eShould I report Cohen’s d or eta squared for repeated measures ANOVA?u003c/strongu003e

Report partial eta squared, and ideally generalized eta squared alongside it. Cohen’s d is built for two-group comparisons, not a repeated measures design with three or more time points.u003cbru003e

u003cstrongu003eCan three separate t-tests replace a repeated measures ANOVA?u003c/strongu003e

No. Running three t-tests on the same participants inflates your familywise error rate to roughly 14%, ignores sphericity completely, and gives no single omnibus test to defend to a committee.u003cbru003e

u003cstrongu003eHow do I report repeated measures ANOVA results in APA style?u003c/strongu003e

State the corrected F-value with its adjusted degrees of freedom, the p-value, and the effect size together in one line, for example: F(1.56, 51.5) = 12.84, p u0026lt; .001, partial η² = .28. That single format covers everything a committee expects to see at once.

Perfect for students, researchers, and professionals looking to build real statistical skills.