Bonferroni vs FDR Correction: Which One Should You Actually Use?
Bonferroni vs FDR correction is one of the most common questions I get from dissertation students the week before their data chapter is due. After 12 years of guiding people through this exact decision, I can tell you most of the confusion is not about the maths. It is about not knowing which one your committee actually wants to see.
The core difference in one line: Bonferroni controls the chance of even a single false positive across all your tests, while FDR controls the proportion of false positives among the results you call significant. That single sentence decides most of what follows in this article.
Use Bonferroni when you are running a small number of planned, confirmatory comparisons and a single false positive would be costly. Use FDR (Benjamini-Hochberg) when you are screening many comparisons and want more statistical power without inflating your error rate too much. Most dissertation-level ANOVA post hoc work with fewer than 10 to 15 pairwise tests is fine with Bonferroni, Holm, or a Tukey style method. Once you cross into dozens or hundreds of tests, such as a survey with many items or a genomics style dataset, FDR is the more sensible choice.
What Is the Multiple Comparison Correction Problem?
Every time you run a hypothesis test at the standard 0.05 alpha level, you accept a 5% chance of a false positive on that one test. Run 20 independent tests and your chance of at least one false “significant” result climbs to roughly 64%. That is not a rare edge case. It happens on almost every dissertation with more than five or six pairwise group comparisons.
This is why a multiple comparison correction exists. It adjusts either your significance threshold or your p-values so that running many tests does not automatically manufacture fake findings. Regulatory bodies such as the FDA and EMA require some form of multiplicity adjustment for confirmatory clinical trials with multiple primary endpoints, which shows how seriously this problem is taken outside academia too.
Family-Wise Error Rate (FWER) vs False Discovery Rate (FDR)
Here is the definition worth remembering, because this is the bit AI answer boxes usually get half right:
Family-wise error rate (FWER) is the probability of making at least one false positive across your whole set of tests. False discovery rate (FDR) is the expected proportion of false positives among only the tests you actually called significant. FWER methods like Bonferroni guard against any mistake. FDR methods like Benjamini-Hochberg accept a controlled rate of mistakes in exchange for finding more true effects.
That difference sounds academic until you are staring at your SPSS output at 1 AM wondering why nothing is significant anymore.
What Is Bonferroni Correction?
Bonferroni correction takes your chosen alpha (usually 0.05) and divides it by the number of comparisons you are running. It is named after the Italian mathematician Carlo Emilio Bonferroni, who developed the underlying inequality in the 1930s, decades before anyone was running post hoc tests on SPSS output.
Formula: α(adjusted) = α ÷ m, where m is the number of tests.
Worked example: Say you ran an ANOVA with four groups, giving you six pairwise post hoc comparisons. Your adjusted alpha becomes 0.05 ÷ 6 = 0.0083. Any comparison needs a p-value below 0.0083 to survive, not the usual 0.05.
I have watched dissertation students run 15 comparisons on a small sample, apply Bonferroni, and see every result that was significant at 0.05 disappear once the correction goes in. That is not a bug. Bonferroni is deliberately conservative, and with a handful of tests on a modest sample it can wipe out genuine effects along with the false ones.
A close relative worth knowing is the Holm correction. Holm still controls the family-wise error rate the same way Bonferroni does. But instead of dividing alpha by the total number of tests every time, it ranks your p-values and applies a progressively less strict threshold to each one. The result is a method that is uniformly more powerful than plain Bonferroni at the same error guarantee, which is why most statisticians now recommend Holm over Bonferroni whenever software allows it.
What Is False Discovery Rate (FDR) / Benjamini-Hochberg Correction?
The FDR approach was introduced by Yoav Benjamini and Yosef Hochberg in their 1995 paper in the Journal of the Royal Statistical Society, and it changed how large-scale testing is done in fields like genomics. Instead of one fixed threshold, it ranks your p-values from smallest to largest and compares each one against a sliding bar.
Procedure (step-up):
- Rank your p-values from smallest (rank 1) to largest (rank m).
- For each rank k, calculate the threshold (k ÷ m) × α.
- Find the largest p-value that is still smaller than its own threshold.
- That p-value, and every p-value smaller than it, is declared significant.
Worked example: With 20 tests and α = 0.05, the p-value ranked 5th needs to beat (5 ÷ 20) × 0.05 = 0.0125, not 0.0025 the way Bonferroni would demand for the same rank. This is why FDR consistently recovers more true findings as the number of tests grows.
A comparative simulation study published in Frontiers in Genetics found that FDR controlling procedures are uniformly more powerful than FWER controlling procedures across the scenarios tested. Several statistics blogs quote that finding and stop there, as if higher power settles the question. It does not.
For a dissertation, power is only half the argument. Your methods section also needs to justify why you accepted a higher tolerance for false positives, and an examiner who studied under the FWER tradition may ask exactly that.
For tests that are negatively or unpredictably correlated with each other, the standard Benjamini-Hochberg method can understate the true false discovery rate. In that situation, the more conservative Benjamini-Yekutieli correction adds a logarithmic penalty to keep the guarantee valid. It is worth mentioning in your methodology if your comparisons are not independent.
Bonferroni vs FDR Correction: Side-by-Side Comparison Table
| Property | Bonferroni | FDR (Benjamini-Hochberg) |
|---|---|---|
| Error controlled | FWER (any false positive) | Proportion of false positives among significant results |
| Statistical power | Lowest | Higher, especially with many tests |
| Procedure type | Single-step, fixed threshold | Step-up, ranked threshold |
| Best for | Small number of confirmatory tests | Large screening or exploratory studies |
| Risk if misused | Misses real effects (Type II error) | Lets through more false positives |
| Common in | Clinical trials, small post hoc ANOVA | Genomics, large surveys, big data screening |
For tests on independent comparisons where you want a touch more power than Bonferroni without moving to FDR at all, the Šidák correction is a smaller, less conservative adjustment that many students overlook.
Before deciding on any correction method, it is also worth confirming your ANOVA assumptions actually hold. A violated normality assumption can matter more to your results than which correction you pick, and the Shapiro-Wilk normality guide walks through how to check this in under five minutes.
Which Correction Should You Use for Your Study?
This is the part most articles skip, and it is the part your supervisor actually cares about.
Use Bonferroni, Holm, or Tukey HSD when:
- You have a small, pre-planned set of comparisons (under 10 to 15).
- A single false positive would misdirect your entire conclusion.
- Your field or journal has a tradition of expecting FWER control, which is common in psychology and management dissertations.
Use FDR when:
- You are running dozens or hundreds of tests, such as a large Likert-scale survey or gene expression data.
- You are in an exploratory phase and want to flag promising effects for follow-up, not make a final confirmatory claim.
- Your discipline already accepts FDR as standard, common in bioinformatics and large-scale psychometrics.
How Correction Affects Your Required Sample Size
A stricter adjusted alpha does not just change which results survive, it changes how much data you need in the first place. Because a smaller critical region demands a larger test statistic to reach significance, a Bonferroni-corrected analysis typically needs a bigger sample than an uncorrected one to detect the same effect at the same power. This is worth checking during your proposal stage, not after data collection. Finding out you are underpowered after the correction has already been applied is a difficult position to write your way out of.
This is a pattern I see often: a student runs a dozen or more pairwise comparisons on a modest sample and applies Bonferroni. Every finding that looked promising at the raw 0.05 level disappears once the correction goes in. In this kind of scenario, switching the justification to a Holm correction, which controls the exact same family-wise error rate with more power, recovers genuine significant pairs without inflating the false positive rate.
The lesson was never to pick whichever method gives better numbers. It was to know, before running the test, which error you are willing to accept, and to be able to defend that choice on paper.
How to Calculate Bonferroni and FDR Correction
- In SPSS: For Bonferroni, go to Analyze > Compare Means > One-Way ANOVA > Post Hoc, and tick “Bonferroni.” SPSS does not have a native FDR option in the post hoc menu, so most students export p-values and analyse them in R instead.
- In R: Use
p.adjust(p_values, method = "bonferroni")orp.adjust(p_values, method = "BH")for Benjamini-Hochberg. One line, both methods, instantly comparable. - In Python: The
statsmodelslibrary has amultipletestsfunction that supports bothbonferroniandfdr_bhmethods, useful if your analysis pipeline is already in Python rather than R. - In Excel: Sort your p-values ascending, add a rank column, then apply the α/m formula for Bonferroni or the (k/m)×α formula for FDR in an adjacent column.
- Using an online tool: If you want to skip manual formulas and check your numbers quickly, our Bonferroni correction calculator gives you the adjusted threshold and a plain-language read on whether your result survives correction. The calculator shows the exact formula behind each result, so you can cite the method itself in your write-up rather than the tool.
If your post hoc analysis involves unequal variances or unequal group sizes, it is also worth comparing against the Games-Howell test calculator or the Tukey HSD calculator, since the correction method you pick should match the post hoc test correction method your design actually calls for.
How to Report Corrected P-Values in APA Format
Report both the raw and adjusted p-value where possible. For example: “The difference between Group A and Group B was significant, p = .012, p(Bonferroni-adjusted) = .048.” For FDR, name the method directly: “p-values were adjusted using the Benjamini-Hochberg procedure to control the false discovery rate at q = .05.” Naming the exact procedure, not just saying “corrected for multiple comparisons,” is what separates a defensible methods section from a vague one.
Common Mistakes When Applying Multiple Comparison Corrections
- Switching methods after seeing the results. Deciding your correction method before you see the p-values is the whole point of controlling error rates. Choosing FDR only because Bonferroni erased your findings is a red flag any careful examiner will spot.
- Redefining your “family” of tests after the fact. If you decide midway to drop three comparisons from the correction because they were not significant anyway, you have quietly inflated your error rate again.
- Treating a corrected p-value as an effect size. A result surviving correction tells you it probably is not noise. It says nothing about how large or meaningful the effect actually is.
Frequently Asked Questions
What is the difference between Bonferroni and FDR correction?
Bonferroni controls the probability of even one false positive across all your tests, while FDR controls the expected proportion of false positives among the results you call significant. Bonferroni is stricter and sacrifices more true findings to guarantee that safety.
u003cstrongu003eIs Bonferroni correction too conservative?u003c/strongu003e
For small numbers of comparisons it is manageable, but past 10 to 15 tests it becomes noticeably strict and can hide real effects, particularly on smaller dissertation samples.
u003cstrongu003eWhat is a good FDR-corrected p-value?u003c/strongu003e
Most researchers work with a false discovery rate of 5% (q = .05), the same convention as the standard alpha level, though exploratory work sometimes uses q = .10 for more sensitivity.
u003cstrongu003eWhich correction is best for ANOVA post hoc tests?u003c/strongu003e
For a handful of pairwise comparisons after ANOVA, Tukey HSD or Bonferroni is standard, and Holm is worth considering wherever your software supports it, since it is more powerful at the same error guarantee.
u003cstrongu003eCan I use FDR correction with a small sample size?u003c/strongu003e
Yes, sample size and number of tests are separate things. FDR works on the number of comparisons, not the number of participants, though very few comparisons make the ranking procedure less meaningful.u003cbru003e
u003cstrongu003eCan I switch from Bonferroni to FDR after running my analysis?u003c/strongu003e
Only if you can justify it methodologically before looking at which specific results changed. Switching purely because Bonferroni erased your findings is a decision most examiners will question.u003cbru003e
u003cstrongu003eWhat is the difference between family-wise error rate and false discovery rate?u003c/strongu003e
FWER controls the chance of even one false positive across all tests. FDR controls the proportion of false positives only among the tests you call significant, which is a more lenient and more powerful standard.
If you are staring at a results chapter and still unsure which correction fits your design, that is exactly the kind of decision worth checking with someone before it becomes a viva question. If you are not even sure whether your analysis needs this level of scrutiny, this piece on knowing when you actually need a dissertation expert is a good place to start. Otherwise, you can go straight to a one-on-one statistics consultation and we can sort out the right method before it becomes a problem in your defence.