Math calculator

Simpson’s Paradox Calculator

When every group says one thing and the total says the opposite.

Does the total contradict the parts

Name, then group 1’s successes and total, then group 2’s. Real data. Men admitted at 44.52% against women’s 30.35% — and women had the higher rate in four of the six departments.

Pooled across all 6 groups

Men 44.52% · Women 30.35%

A gap of 14.1645 percentage points favouring Men. Within the groups, 2 favour Men and 4 favour Women.

Simpson’s paradox: the total says Men and 4 of the 6 decided groups say Women. Neither figure is a mistake. They answer different questions, and the pooled one silently answers a question about where people applied rather than how they were treated.
Rate for each group within each stratum, with the stratum’s own overall rate and its composition
GroupMenWomenHigherOverall rateShare Men
Dept A62.06%82.41%Women64.42%88.42%
Dept B63.04%68.00%Women63.25%95.73%
Dept C36.92%34.06%Men35.08%35.40%
Dept D33.09%34.93%Women33.96%52.65%
Dept E27.75%23.92%Men25.17%32.71%
Dept F5.90%7.04%Women6.44%52.24%
Holding composition constant: Men 38.73% against Women 43.00%. That is direct standardisation — apply both groups’ own within-group rates to one shared set of weights, so the mix can no longer drive the comparison. The gap moves from +14.1645 points to -4.2637 points — and changes sign, which is the paradox resolved rather than restated. This is the actual repair, not a caveat: it is what a rate comparison should have been all along whenever the groups are distributed differently across strata.

Men pooled

44.52%

2691 total

Women pooled

30.35%

1835 total

Spread in group rates

57.9733 pp

how much the strata differ

Spread in composition

63.0210 pp

how unevenly they mix

Both spreads are large here — 57.973points between the strata’s own rates, and 63.021 points in how the two groups are mixed across them. That combination is what makes a reversal possible. If both groups had the same rate in every stratum there would be no direction to contradict. If the composition were identical everywhere, the pooled difference would be a positively-weighted average of the stratum differences and could not overturn a unanimous verdict — though a bare majority can still be outweighed by one large stratum. This is why a randomised experiment is immune: randomisation equalises the composition by construction, while observational data produces both spreads by the same mechanism that created the groups.

What this tool shows

Berkeley, 1973: men admitted at 44.52%, women at 30.35%. A 14-point gap — and women had the higher rate in four of the six largest departments. This detects that reversal in your own data and then repairs it, which takes the Berkeley gap to −4.26 points.

  • Rates for both groups within every stratum
  • The pooled rates, and whether they contradict the strata
  • Automatic detection of a reversal, and of a total reversal
  • Direct standardisation — the actual repair
  • How much the strata differ, and how unevenly they mix
  • Two real published datasets built in
Reversal detected Real Berkeley data Standardisation included Both spreads reported

Standardising flips the Berkeley gap from +14.16 to −4.26 points.

Updated 8 September 2026 · Works in any browser, no installation

Simpson’s paradox is when a comparison reverses on pooling. Group A beats group B in every subgroup, and loses overall. Nothing has gone wrong with the arithmetic — both figures are correct, and they answer different questions. The paradox is that only one of those questions was asked out loud.

At a glance

Formula shown
A pooled rate is a weighted average of stratum rates, with weights set by where the group’s members happen to be. Reversal needs two things at once: the strata must differ in their base rates, and the groups must be distributed differently across them. Direct standardisation replaces each group’s own weights with a shared set, so composition can no longer drive the comparison.
Scenario support
Comparing admission, hiring or approval rates between groups; treatment success rates across patient severities; conversion rates across traffic sources; any rate comparison where the two groups are not spread evenly across the categories.
Educational estimate
Planning support from the values you enter — not professional advice.

Berkeley, 1973 — the case that made it famous

The tool loads real data by default, not a constructed example. These are UC Berkeley’s graduate admissions figures for the six largest departments in autumn 1973, as published in Science in 1975.

Pooled, men were admitted at 44.52% and women at 30.35% — a gap of 14.16 percentage points, and on its face a serious finding.

Department by department, women had the higher admission rate in four of the six. In Department A, 82.4% of women were admitted against 62.1% of men. In Department B, 68.0% against 63.0%.

The explanation is where people applied, not how they were treated. Departments A and B admitted around 60% of applicants and received 1,385 male applications against 133 female. Department F admitted about 6% and received 341 female applications against 373 male. Women applied in far greater numbers to the departments that admitted almost nobody, so their pooled rate was dragged toward those departments’ base rates.

The original paper’s conclusion is the part usually lost. Bickel and colleagues did not conclude “no bias”. They concluded that the bias was not in the admissions decisions, and pointed at what puts women disproportionately into the most competitive fields in the first place. The paradox relocated the question rather than dismissing it.

The tool’s second preset is the other classic, from a 1986 comparison of kidney stone treatments. Treatment A wins on small stones (93.1% against 86.7%) and on large stones (73.0% against 68.8%), and loses overall (78.0% against 82.6%). That is a total reversal — every stratum one way, the pooled figure the other — because the more effective treatment was given to the harder cases.

It needs two things to be true at once

The reversal is not a fluke of small samples, and it is not a paradox in the logical sense. It is a property of weighted averages, and it requires exactly two conditions.

First, the strata must differ in their base rates. If every department admitted 40% of applicants, no distribution of applicants could move a pooled rate anywhere. The tool reports this spread — at Berkeley it is 58 percentage points between the easiest and hardest department.

Second, the two groups must be distributed differently across those strata. If men and women had applied to each department in the same proportions, pooling would be a fair weighted average with identical weights. The tool reports this spread too — at Berkeley it is 63 points.

The two conditions are not symmetric, and it is worth being exact. If both groups have the same rate in every stratum, no stratum has a direction and there is nothing to reverse. If the composition is identical across strata, the pooled difference is a positively-weighted average of the stratum differences, so it cannot contradict a unanimous verdict — a total reversal becomes impossible. A bare majority can still be outvoted by a single large stratum, so equal composition does not rule out every reversal, only the complete kind. The verification suite asserts exactly that distinction, because an earlier draft of this page claimed the stronger version.

This is why randomised experiments are immune. Randomisation forces the second condition to fail: the groups are distributed across every stratum in the same expected proportions, so composition cannot differ systematically. It is one of the sharpest arguments for randomisation there is, and it does not depend on any assumption about the outcome.

And it is why observational data is so exposed. In real data people sort themselves into strata for reasons connected to the outcome — sicker patients get the stronger treatment, harder applicants apply to harder departments — which produces both conditions at once, by the same mechanism.

Standardisation is the repair, not a caveat

Most treatments of this paradox stop at “be careful with pooled rates”. There is an actual calculation that fixes it, and the tool performs it.

Direct standardisation applies both groups’ own within-stratum rates to one shared set of weights. Each group keeps the rates it actually achieved; only the mix is equalised. The result answers “what would each group’s overall rate have been if they had been distributed the same way?” — which is the question a fairness comparison was always trying to ask.

On Berkeley it changes the answer. The raw gap is +14.16 points favouring men. After standardising it is −4.26 points, favouring women. An 18-point swing, and a change of sign, from nothing but the weighting.

The choice of weights is a real decision, and it should be stated. This tool uses the combined totals, which is the conventional choice when neither group is privileged as the standard. Using one group’s distribution instead gives a different number and answers a different question — “what would women’s rate be if they applied like men?” is not the same as “what would both be under a common standard?”

It only controls for the strata you have. Standardising by department removes department composition and nothing else. If another variable is doing the same work — and you did not measure it — the standardised figure is exposed to it exactly as the raw one was. This is the same limit that partial correlation runs into.

Which number is right depends on the question

The temptation is to declare the stratified figures “true” and the pooled one “misleading”. That is usually right and it is not automatic, and the difference matters.

Ask about the process and the strata win. “Does this department treat men and women differently?” is answered department by department. Pooling contaminates it with where people applied, which the admissions committee did not choose.

Ask about the outcome and the pooled figure is a fact. “What share of women who applied to Berkeley got in?” is 30.35%. That is not misleading; it is the answer. It matters for capacity planning and for describing what actually happened to a cohort.

The genuine rule is causal, not statistical. Condition on a variable when it is a common cause of both the group and the outcome. Do not condition on it when it sits on the causal path between them — if the field someone applies to is itself shaped by the discrimination you are measuring, then adjusting for department removes part of the very effect you are looking for. No amount of data resolves this; it takes an argument about how the world works.

Which is why the honest presentation is both. Report the stratum rates, the pooled rates and the standardised comparison together, and say what each answers. A single number here is a choice about the question, and presenting it alone hides that a choice was made.

Where it turns up outside the textbooks

Four settings where this is a live risk rather than a curiosity, and the strata that usually cause it.

Conversion rates by traffic source. A new channel can convert better on every device and worse overall, because it brings disproportionately more mobile traffic and mobile converts worse everywhere. Stratify by device before comparing channels.

Hospital and surgeon outcomes. The best surgeons take the hardest cases, so raw mortality rates rank them backwards. Every published surgical outcome table is risk-adjusted for exactly this reason, and unadjusted league tables actively punish the specialists.

Batting averages across seasons. A player can hit for a higher average than another in both of two seasons and a lower combined average, if their at-bat counts differ between a good year and a bad one. It has genuinely happened in Major League Baseball.

Pay gap analyses. An organisation can pay women more than men in every job grade and still report a large aggregate gap, because grade composition differs. Both numbers are real and they support opposite headlines — which is why the reporting standard asks for both.

The common signature is a comparison across groups that are not distributed alike. Whenever you are about to compare two rates, ask what the two populations differ in besides the thing you are measuring. If you can name it, stratify by it before you compare.

Related reversals that are not this one

Three effects get filed under this name and are different mechanisms with different fixes.

The ecological fallacy is about level, not reversal. A correlation between group averages need not hold within groups — and can even run the opposite way. Country-level income and some outcome can correlate strongly across 40 nations and be near zero inside each. Simpson’s paradox is one route to this, but the fallacy is the broader error of transferring a group-level finding to individuals.

Berkson’s paradox is created by selection. Two independent traits become negatively correlated among people selected for having at least one. Among hospital patients, two unrelated conditions appear to protect against each other; among admitted students, test scores and interview scores often correlate negatively. Stratifying does not fix this — the sample itself is the problem.

Regression to the mean is about repetition. An extreme measurement is followed by a less extreme one for reasons of noise alone, which makes any intervention applied to the extreme group look effective. No stratification helps; only a control group does.

What distinguishes Simpson’s paradox is that all the data is present and correct. Nothing was selected away, nothing was measured twice, and every rate in the table is accurate. The reversal comes purely from how the parts were weighted into the whole — which is why it can be repaired arithmetically, and the other three cannot.

Sources and methodology

References for both datasets and the adjustment.

Method. Every stratum’s two rates and the two pooled rates are computed from the raw counts, so the reversal is detected rather than eyeballed. A reversal is reported only when a strict majority of the strata that have a direction contradict the pooled direction — a three-three split is not a reversal — and a total reversal is reported separately when every stratum disagrees with the total. Direct standardisation uses the combined stratum totals as weights, which is the conventional neutral choice, and the page says so because a different weight set answers a different question. The suite asserts the Berkeley figures to four decimals (44.5188% against 30.3542%, women higher in four of six departments), that standardising moves the gap from +14.16 to −4.26 percentage points and changes its sign, that the kidney stone data is a total reversal, and — checked by construction over generated data rather than by example — that exactly equal stratum rates make any reversal impossible while equal composition rules out only the total kind, since a bare majority can be outweighed by one large stratum. That engine is verified on every change against 40 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

ProbabilityTwo events, repeated trials and Bayes, with the three usual errors handled — the dropped overlap in P(A or B), n×p instead of the complement, and the base rate that makes a 99% test 17% right.
Correlation CoefficientReports Pearson, Spearman and Kendall together with the scatter plot, and ships Anscombe's quartet built in — four datasets with an identical r of 0.816 that Spearman tells apart.
Odds RatioOdds ratio, relative risk, risk difference and number needed to treat from one 2x2 table — because an odds ratio of 6.00 can describe a relative risk of 1.50.
PercentageSolve X% of Y, what percent X is of Y, reverse percentage, increase/decrease, discounts, and tax, tip, or commission.
Weighted AverageEach value carries the weight you give it, with every item's share of the total shown as a percentage so you can see what is actually driving the answer.
Chi-SquareGoodness of fit and tests of independence with every expected count and per-cell contribution shown — because the validity condition is about expected counts, not observed ones, and most calculators hide them.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool. Standardisation removes only the composition you stratify by; an unmeasured variable acting the same way is untouched. Whether to adjust for a variable at all is a causal question, not a statistical one — adjusting for something on the causal path removes part of the effect being measured.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (3 updates)

Published 8 September 2026

  1. Published a Simpson's paradox tool that detects the reversal rather than describing it, built on real published data: UC Berkeley's 1973 graduate admissions, where men were admitted at 44.52% against women's 30.35% while women had the higher rate in four of the six largest departments.
  2. Standardises as well as detects, which almost every treatment of this paradox omits. Applying both groups' own within-stratum rates to shared weights takes the Berkeley gap from +14.16 percentage points to -4.26 — a change of sign, an 18-point swing from the weighting alone.
  3. Corrected a claim in an earlier draft. I had written that a reversal is impossible when either the base-rate spread or the composition spread is zero; only the first half holds. With identical composition the pooled difference is a positively-weighted average of the stratum differences, so it cannot overturn a unanimous verdict, but a bare majority can still be outweighed by one large stratum. Both directions are now asserted over a thousand generated datasets.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.