Kruskal Wallis Test Calculator
Kruskal-Wallis Test
Computation Complete
Kruskal-Wallis Results
Step-by-Step Calculation
Interpretation
Chi-Square Distribution & Critical Rejection Region
I have reviewed close to five hundred dissertations and research projects over the last twelve years. The Kruskal Wallis test is one of the tests students run correctly but write up wrong, more often than any other test on this list. This guide, and the Kruskal-Wallis test calculator built into it, fixes that. Use it right away, then work through the full hand-calculation walkthrough so you understand exactly what the tool is doing behind the screen.
By the end of this page you will be able to run the test, interpret the H statistic correctly, follow it up with a proper post-hoc analysis, and write a results section your supervisor will not send back for corrections.
What Is the Kruskal Wallis Test? My Quick Take After 12+ Years of Dissertation Reviews
The Kruskal-Wallis test is a nonparametric method that compares three or more independent groups by ranking all observations together and checking whether the rank distributions differ across groups. It replaces means and variances with ranks, which is exactly why it survives skewed data, outliers, and ordinal scales that would break a parametric test.
Most textbook definitions stop right there and jump straight to the formula. In my experience that is where students lose the plot, because the formula means nothing until you actually know when you are supposed to reach for it.
You need the Kruskal-Wallis test for non-normal data the moment your dataset fails a normality check. It also fits Likert-scale responses collected across departments, treatment arms, or product variants, where the numbers are ordinal rather than truly continuous. It is purpose-built for a Kruskal-Wallis test for three or more groups scenario.
If you only have two groups, you do not need this test at all. A Mann-Whitney U test does the same job with less machinery.
Assumptions You Still Cannot Skip
Being nonparametric does not mean assumption-free. You still need:
- Independent observations, no participant appears in more than one group
- Groups drawn from populations with a similar shape of distribution, even if that shape is not normal, checked upfront with the Shapiro-Wilk test for normality or Levene’s test for equal variance
- A dependent variable measured on at least an ordinal scale
Skip the independence check and no amount of rank-based cleverness will save your result.
Null Hypothesis Explained
The Kruskal-Wallis test null hypothesis states that all groups come from populations with the same distribution, which in practice means the population medians are equal. The alternative hypothesis says at least one group’s distribution differs from the rest.
Here is a mistake I see constantly in draft chapters. Students write the null as “the means are equal.” That is an ANOVA sentence, not a Kruskal-Wallis sentence. This test never touches means, so keep your wording rank and distribution based, and your examiner will notice the precision.
One more mix-up worth clearing up early: this is not the chi-square test in disguise. Chi-square tests whether two categorical variables are associated. Kruskal-Wallis compares the rank distribution of a continuous or ordinal variable across three or more groups. Different question, different test, even though both eventually lean on the chi-square distribution to get a p-value.
Try the Free Kruskal Wallis Test Calculator
This is a free Kruskal-Wallis test calculator online, no login required and no paywall on the core output. Paste your group data, choose your significance level, and you get the H statistic, degrees of freedom, p-value, effect size and a full post-hoc breakdown on one screen.
Most free calculators I tested while researching this page stop at H and a p-value. That is not a Kruskal-Wallis test calculator with steps. It is just a number generator. A p-value alone tells your reader almost nothing about which groups actually differ or how large the effect is, and that gap is exactly why so many results chapters get a “so what” comment from a supervisor.
What You Need Before You Enter Data
- Data from three or more independent groups, not paired or repeated measures
- Group values, one column or list per group
- A chosen significance level, usually 0.05
- Group labels, so the post-hoc output reads clearly instead of “Group 1 vs Group 2”
What the Calculator Gives You
- H statistic and degrees of freedom
- Exact p-value, not just “significant” or “not significant”
- Epsilon-squared effect size
- Dunn’s post-hoc pairwise comparisons with adjusted p-values
How to Calculate the Kruskal-Wallis H Statistic by Hand
This section covers how to calculate the Kruskal-Wallis H statistic step by step, with a worked example using real numbers. If you need to calculate the Kruskal-Wallis test by hand for an exam or a viva, follow these four steps. I still make my mentees do this by hand at least once, even though they will use software for the real analysis. It is the fastest way to actually understand what the calculator is doing, and it is exactly what most examiners ask you to reproduce in a viva.
Here is the formula:
H = [12 / (N(N+1))] × Σ(Rᵢ² / nᵢ) − 3(N+1)
Where N is the total sample size across all groups, k is the number of groups, Rᵢ is the sum of ranks in group i, and nᵢ is the sample size of group i.
The steps, in order:
- Rank every observation from every group together, smallest to largest
- Sum the ranks separately for each group
- Plug the rank sums into the H formula
- Compare H against the chi-square distribution at k−1 degrees of freedom
Step-by-Step Worked Example
Say I am comparing exam scores from three teaching methods for a client’s internal training study.
- Method A (n=5): 45, 52, 38, 61, 49
- Method B (n=6): 55, 62, 58, 70, 65, 60
- Method C (n=4): 30, 35, 42, 28
Step 1: Rank every value together.
With 15 total observations and no tied scores in this dataset, ranks run cleanly from 1 to 15.
If two scores are identical, each tied value gets the average of the ranks it would have occupied. This tie correction is something most calculators apply automatically and most manual write-ups skip entirely, which quietly inflates the H statistic. The correction factor is C = 1 − [Σ(t³−t) / (N³−N)], where t is the number of values tied at each rank, and the corrected H = H ÷ C.
Step 2: Sum the ranks per group.
- Method A rank sum (R₁) = 37
- Method B rank sum (R₂) = 72
- Method C rank sum (R₃) = 11
Check: 37 + 72 + 11 = 120, which equals N(N+1)/2 = 15×16/2 = 120. If your totals do not match this check, you have a ranking error somewhere and need to redo it before going further.
Step 3: Apply the formula.
Rᵢ²/nᵢ values: 37²/5 = 273.8, 72²/6 = 864, 11²/4 = 30.25. Sum = 1168.05.
H = [12/(15×16)] × 1168.05 − 3(16) = (0.05 × 1168.05) − 48 = 58.4 − 48 = 10.40
Step 4: Compare H to the chi-square distribution.
Degrees of freedom = k − 1 = 2. At df=2 and alpha=0.05, the critical chi-square value is 5.99. Our H of 10.40 clears it comfortably, and the exact p-value works out to approximately 0.006.
We reject the null hypothesis. At least one teaching method produces a score distribution that differs from the others.
Running the Kruskal-Wallis Test in SPSS, R, Python and Excel
The calculator on this page is the fastest route, but you will still be asked which software you used, so here is the one-line version for each.
- SPSS: Analyze > Nonparametric Tests > Legacy Dialogs > K Independent Samples, select Kruskal-Wallis H, and SPSS returns the chi-square value, df and asymptotic significance directly.
- R:
kruskal.test(score ~ group, data = df)gives you the H statistic (labelled chi-squared), degrees of freedom, and p-value in one line. - Python:
from scipy.stats import kruskalthenkruskal(group_a, group_b, group_c)returns the statistic and p-value as a named tuple. - Excel: there is no native Kruskal-Wallis function. You have to rank every value yourself with
RANK.AVG(), sum ranks per group manually, and build the H formula in a separate cell. This is precisely why most people searching for a Kruskal-Wallis test calculator with steps end up on a page like this one instead of fighting Excel formulas.
Kruskal-Wallis Test vs ANOVA: Which One Should You Use?
Knowing when to use Kruskal-Wallis vs ANOVA comes down to one question: does your data meet the normality assumption or not? This is the single most searched comparison in this entire topic, and most articles answer it with a bullet list of assumptions and nothing about the actual decision process. Here is how I decide it for a client’s dataset in practice.
| Aspect | One-Way ANOVA | Kruskal-Wallis Test |
|---|---|---|
| Data requirement | Continuous, normally distributed | Ordinal or non-normal continuous |
| What it compares | Group means | Group rank distributions |
| Key assumption | Normality and equal variance | Independence only, no normality needed |
| Test statistic | F | H, approximated by chi-square |
| Sensitivity to outliers | High | Low, ranks absorb extreme values |
| Post-hoc test | Tukey HSD, Games-Howell | Dunn’s test, Conover-Iman |
My practical rule: run the one-way ANOVA calculator only after confirming normality. If your Shapiro-Wilk p-value comes back under 0.05, or your groups are small and visibly skewed, switch to Kruskal-Wallis instead of forcing ANOVA through and hoping nobody checks your residual plots. Still unsure which test your specific design calls for? Run it through the test-selection guide before you commit to either one.
Interpreting Your Results: A Full Example
A significant H statistic only tells you that the groups differ somewhere. It does not tell you which groups, and it does not tell you how much they differ by. Both need a second step, which brings us to the Kruskal-Wallis test interpretation example most guides skip entirely.
From our worked example above: H(2, N=15) = 10.40, p = .006. This is statistically significant at alpha = 0.05, so we reject the null hypothesis that all three teaching methods produce equal score distributions.
If Your Result Isn’t Significant
A p-value above 0.05 does not prove the groups are equal. It only means you did not find enough evidence of a difference. Small per-group samples make it easy to miss a real effect, so before you write “no significant difference” in your conclusion, check your sample size honestly.
Report the non-significant result exactly as it is, with the same H, df and p-value format used for a significant one. Do not go hunting for a different test until you get the answer you wanted. That is the fastest way to fail a viva.
Calculating Effect Size
A p-value never tells you if a difference actually matters in practice, only whether it is unlikely to be chance. Epsilon-squared closes that gap.
ε² = H / (N − 1)
For our example: ε² = 10.40 / 14 = 0.74, a large effect by conventional benchmarks, where anything above 0.14 is generally treated as large. A close cousin, eta-squared-H, uses η²H = (H − k + 1) / (N − k) and usually lands on a very similar verdict. Use whichever your supervisor or target journal expects to see.
Post-Hoc Analysis: Which Groups Actually Differ?
Here is the step most dissertation drafts get wrong. A significant Kruskal-Wallis result is an invitation to look closer, not a conclusion by itself. You still need a Kruskal-Wallis test with post-hoc analysis to say which specific pairs differ.
Dunn’s test is the standard choice. It compares mean ranks between every pair of groups and adjusts the p-values for multiple comparisons, usually with a Bonferroni or Holm correction.
Worked Pairwise Example: Method B vs Method C
Mean rank of B = 72/6 = 12. Mean rank of C = 11/4 = 2.75. Difference = 9.25.
z = (12 − 2.75) / √[(15×16/12) × (1/6 + 1/4)] = 9.25 / √(20 × 0.4167) = 9.25 / 2.89 = 3.20
With a Bonferroni-adjusted alpha for three comparisons (0.05/3 ≈ 0.0167), the critical z value is roughly 2.39. Our z of 3.20 clears that easily, so Method B differs significantly from Method C. The calculator on this page runs this for every pair automatically. You do not need to do this by hand for a real submission, but you should know what it is doing under the hood.
Writing Up Kruskal-Wallis Results in APA Style
Copy this template and swap in your own numbers:
“A Kruskal-Wallis H test showed a statistically significant difference in exam scores across the three teaching methods, H(2, N = 15) = 10.40, p = .006, ε² = .74. Post-hoc pairwise comparisons using Dunn’s test with Bonferroni correction indicated that Method B scored significantly higher than Method C (p < .05).”
Notice this sentence reports four separate things: the test statistic with degrees of freedom, the sample size, the exact p-value, and the effect size. Drop any one of these and a strict reviewer will ask for it back before approving your chapter.
Common Mistakes I See in Dissertations Using This Test
- Reporting means instead of medians or rank sums. This test never compares means, so do not write “Group A had a higher mean” in your results section.
- Skipping the post-hoc step entirely. A significant omnibus result with no follow-up is an incomplete analysis, not a finished one, and examiners catch this immediately.
- Ignoring tied ranks. Likert-scale data is full of ties, and skipping the correction inflates your H statistic more than most students realise.
- Using it on paired or repeated-measures data. If the same participants were measured under three conditions, you need a Friedman test, not this one.
- Running it with tiny group sizes. Below about five observations per group, the chi-square approximation gets shaky. Flag this limitation honestly in your methodology section instead of hiding it.
A Short Case From My Consulting Work
This is a composite drawn from a pattern I see constantly across client engagements, with details changed for confidentiality. A retail management student had customer satisfaction scores, on a 1 to 5 scale, from three store locations. Normality was never on the table with data that discrete. She ran a Kruskal-Wallis test and got H(2, N=142) = 14.2, p = .0008, then panicked because her supervisor asked “which store, exactly.”
We ran Dunn’s post-hoc test together, found her flagship store differed significantly from the other two, and rewrote her results section using the exact template above. Her viva had zero follow-up questions on that section, which almost never happens with a nonparametric result in my experience.
FAQ
Can Kruskal-Wallis be used for only two groups?
Technically yes, the math still runs, but it becomes mathematically identical to a Mann-Whitney U test at that point. Use a Mann-Whitney U test directly for two groups. It is built for exactly that comparison and saves you a step.
What post-hoc test should I use after Kruskal-Wallis?
Dunn’s test with Bonferroni or Holm correction is the standard choice, and it is what most journals expect to see cited. Conover-Iman is a slightly more powerful alternative if you want to justify a less conservative approach in your methodology.
What is the minimum sample size per group?
There is no hard cutoff, but I get uneasy below five observations per group. Small groups make the chi-square approximation unreliable, so plan ahead with the sample size calculator before you collect data, and note the limitation honestly if your data forces your hand anyway.
Is Kruskal-Wallis a parametric or nonparametric test?
Nonparametric. It makes no assumption about the shape of the underlying population distribution, which is exactly why it works on ranks instead of raw values.
How do I report Kruskal-Wallis results in APA format?
Report the H statistic with degrees of freedom and sample size, the exact p-value, and an effect size, in that order: H(df, N) = value, p = value, ε² = value. Follow it with your post-hoc findings in a separate sentence.
Kruskal and Wallis first introduced this test in a 1952 paper in the Journal of the American Statistical Association, and the core method has barely changed since. That staying power is the best argument I can give you for trusting it in your own methodology chapter.
Siddharth Gupta, dissertation statistics consultant with 20+ years across analytics and research support. Connect on LinkedIn