Before and after, on the same subjects — where only the changers matter.
Only the pairs that changed count
n = 50. chi-square 9.00, exact p 0.0041.
n = 50, of whom 25 changed and 25 did not
p = 0.002700
chi-square = 9.00000 on 1 degree of freedom. The exact binomial p is 0.004077.
chi-square
9.000000
(b − c)²/(b + c)
p (chi-square)
0.002700
Exact binomial p
0.004077
no approximation
With continuity correction
0.005110
Edwards
Changed yes → no
20
40.00% of subjects
Changed no → yes
5
10.00% of subjects
Did not change
25
contributes nothing
Odds ratio for change
4.0000
b / c
25 subjects did not change, and they appear nowhere in the statistic. chi-square is (20 − 5)²/(20 + 5) = 9.00000. No a and no d. Load the “huge study” preset: 1,425 subjects instead of 50, and the identical chi-square, p-value and exact p. That is correct rather than a defect — a subject who answered the same way twice tells you nothing about which direction change tends to run in — and it is the reason a paired design with few changers is underpowered however many subjects it recruited.
Edwards’s continuity correction gives 7.84000 and p = 0.005110. It subtracts 1 from |b − c| before squaring, so it always lowers the statistic and always raises the p-value. It was introduced to bring the chi-square closer to the exact binomial, and it over-corrects in much the same way Yates’s does on a 2×2 independence table. With the exact test one click away there is little reason to use either approximation.
The odds ratio for change is 4.0000 — 20 moved one way against 5 the other. That is the effect size, and it is the thing the p-value does not tell you. Note that it is an odds ratio for changing, conditional on having changed at all, not a ratio of the two marginal proportions — which is a different quantity and the one people usually quote by mistake. The marginal proportions here moved from 60.00% to 30.00%.
What this tool shows
The subjects who did not change contribute nothing to the statistic. chi-square is (b − c)²/(b + c), with no a and no d in it — so the built-in presets with n = 50 and n = 1,425 give the identical chi-square, p-value and exact p. That is right, and it is why a paired study with few changers is underpowered however many people it recruited.
chi-square on 1 degree of freedom, and the exact binomial p beside it
Edwards’s continuity correction, with what it costs
Which p-value to report when there are fewer than 25 discordant pairs
The odds ratio for change, and why it is not the ratio of the marginals
The marginal proportions before and after, computed separately
A demonstration that the concordant cells cancel out entirely
Paired binary Concordant cells ignored Exact p included Odds ratio for change
Updated 12 September 2026 · Works in any browser, no installation
McNemar’s test asks whether a paired yes/no measurement changed. Same subjects, before and after — or two tests applied to the same cases. It is the paired counterpart to a chi-square test of independence, and using that instead is the most common error, because it treats the two measurements as coming from different people.
At a glance
Formula shown
With the paired table [[a, b], [c, d]] where b and c are the discordant counts, χ² = (b − c)²/(b + c) on 1 degree of freedom. Edwards’s correction uses (|b − c| − 1)² instead. The exact version is a two-sided binomial test of b successes in b + c trials at p = 0.5, which is what the chi-square approximates and what should be used below about 25 discordant pairs.
Scenario support
Before-and-after on the same people, two diagnostic tests on the same samples, matched case-control pairs, an A/B test where each user sees both variants, and any binary outcome measured twice on the same subject.
Educational estimate
Planning support from the values you enter — not professional advice.
Why the sample size barely matters
This is the property that makes the test look broken the first time, and it is exactly right.
The statistic contains only b and c. (b − c)²/(b + c). The subjects who answered the same way twice are not in it anywhere.
So two studies with the same discordant counts give the same answer. The presets show n = 50 and n = 1,425 returning an identical chi-square of 9.00, an identical p of 0.0027 and an identical exact p of 0.0041.
The unchanged subjects carry no information about direction. Someone who said yes both times has told you nothing about whether change runs from yes to no or the other way, which is the only question the test asks.
Which means the power depends on b + c, not on n. A study of ten thousand people in which eight changed has the power of a study of eight people. Recruiting more subjects helps only insofar as it produces more changers.
And it argues for choosing subjects who might change. A paired design testing an intervention on a population where almost nobody would have moved either way is expensive and uninformative, and the sample size is not what fixes it.
The exact test, and when the approximation fails
The chi-square version is the one in textbooks, and it is an approximation to something simpler that has no approximation in it.
Under the null, each discordant pair is a coin flip. If change is equally likely in both directions then b is binomial on b + c trials at p = 0.5. That is the whole test.
So the exact p is a two-sided binomial p, and it needs no large-sample assumption at any size. The chi-square approximates it and stops being reliable below about 25 discordant pairs.
The tool reports the exact figure automatically in that case, and prints both either way. On the few-discordant-pairs preset the two differ enough to matter.
Edwards’s continuity correction is the middle option and over-corrects. It subtracts 1 from |b − c| before squaring, always lowering the statistic and raising the p-value, in much the way Yates’s correction does on an independence table. With the exact test available there is little reason for either.
A binomial test is also what you extend for unequal designs. If the null is that change runs one way with some probability other than a half — which happens in matched designs with unequal matching ratios — the binomial takes that p directly, and the chi-square form does not generalise.
Not a chi-square test of independence
Both work on a 2×2 table and they answer different questions on different data. Using the wrong one is the most common error here.
A chi-square test of independence needs independent observations. Its table counts different people in each cell, and it asks whether two attributes are associated.
McNemar’s table counts the same person twice. Each cell is a pair of answers from one subject, and the question is whether the marginal proportions moved.
Applying the independence test to paired data ignores the pairing, which is the whole design. It throws away the matching that made the study efficient, and it gets the standard error wrong in a direction that depends on the correlation between the two measurements.
The tables also look different. An independence table has one row per group; a McNemar table has one row per first-measurement answer. A table whose rows and columns are the same two categories is usually a paired one.
And the pairing is usually worth having. Because each subject is their own control, a paired design removes between-subject variation entirely — the same reason the paired rank test beats its two-sample counterpart, and the reason the power calculator treats paired designs separately.
The effect size, and the one people quote instead
A p-value says whether the change is detectable. Two different numbers describe how large it is, and they are not interchangeable.
The odds ratio for change is b/c. Among subjects who changed at all, how many times more likely was one direction than the other. It is conditional on having changed.
The marginal difference is a different quantity. The proportion answering yes before, minus the proportion answering yes after. That is what a reader usually means by “the effect”, and the tool prints both.
They can tell quite different stories. A large odds ratio among a handful of changers is a small marginal shift; a modest odds ratio among many changers can be a large one.
The confidence interval belongs on the marginal difference. It is the quantity with an interpretation outside the study, and for a paired proportion difference the interval has to account for the pairing — an unpaired interval is too wide.
And significance is not size here either. With enough discordant pairs a 51-to-49 split is significant, and the marginal shift it corresponds to may be a fraction of a percentage point. Report both, as everywhere else.
Where it is the right test
Four situations produce a McNemar table, and two of them are not obviously “before and after”.
The same subjects measured twice. An opinion before and after a campaign, a symptom before and after treatment, a behaviour before and after a policy change.
Two tests applied to the same samples. Comparing two diagnostic assays on one set of specimens: the pairs are the specimens, and the question is whether the two tests differ in how often they call positive.
Matched case-control pairs. Each case matched to a control on age and sex; the pairs are the matched sets, and the discordant ones are where exposure differed. This is the classic epidemiological use and the reason the test is in every textbook on the subject.
Within-subject A/B tests. Where every user sees both variants, the pairing is the user, and treating the two arms as independent groups both loses power and overstates the standard error.
For more than two categories, use Bowker’s or Stuart-Maxwell. Bowker’s test generalises McNemar to a square table of any size, testing symmetry; Stuart-Maxwell tests marginal homogeneity, which is a slightly different hypothesis. Neither is this tool.
And for more than two time points, use Cochran’s Q. It is the extension of McNemar to several repeated binary measurements, in the way ANOVA extends the t test.
What it assumes, and what it does not
McNemar asks very little, which is why it survives on small and awkward data. What it does ask is worth stating.
The pairs must be independent of each other. Within a pair the two measurements are deliberately dependent — that is the design — but one subject’s pair must not influence another’s. Clustered pairs break this and nothing else in the test notices.
It assumes nothing about the distribution. The outcome is binary, so there is no shape to assume, and the null distribution of the discordant counts is exactly binomial.
It tests marginal homogeneity, not agreement. Two raters can agree badly and still have identical marginals, which McNemar will not flag. Agreement is what Cohen’s kappa measures, and the two answer opposite questions about the same table.
It says nothing about which measurement is right. If one of the two is a gold standard, the question is accuracy rather than symmetry, and sensitivity and specificity are the quantities — with the Bayes calculator turning those into what a positive means.
And the direction of change is only as meaningful as the design. A significant result says the marginal proportion moved, not that the intervention moved it. Regression to the mean and simple time trends produce the same table, and only a control group separates them.
Method. The exact p is a two-sided binomial test on the discordant pairs at p = 0.5, computed from the binomial CDF rather than approximated, and it is the figure the tool highlights below 25 discordant pairs. Chi-square and Edwards’s corrected version are reported alongside so the size of each approximation is visible. The suite asserts the property the page rests on by requiring three tables with n = 50, 1,425 and 25 to produce byte-identical chi-square, p-value and exact p; it checks the exact p against a direct factorial computation of the binomial tail, that swapping b and c leaves the two-sided result unchanged across 500 random tables, and that the continuity correction never raises the statistic. That engine is verified on every change against 50 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Chi-SquareGoodness of fit and tests of independence with every expected count and per-cell contribution shown — because the validity condition is about expected counts, not observed ones, and most calculators hide them.
Fisher's Exact TestThe exact p for a 2×2 table under all three two-sided conventions, because they disagree — 0.0406 against 0.0699 on the built-in table, across the 5% line — plus the test's actual size by enumeration, which is 2.30% at a nominal 5%.
Cohen's KappaKappa with the two figures that explain it: the maximum the marginals permit, and PABAK. Two built-in tables with identical 85% agreement give kappas of 0.6995 and 0.3219, and a third with 94.4% agreement gives −0.0234.
Wilcoxon Signed-RankReports how many zero differences it dropped and gives the Hodges-Lehmann shift, because the classical and Pratt variants disagree on the same data and the median of the differences is not what this test estimates.
Statistical PowerPower and sample size from the non-central t rather than a normal approximation, with the gap shown — plus a live demonstration that post-hoc power is a function of the p-value alone, and 0.500044 at p = 0.05 for every study ever run.
Odds RatioOdds ratio, relative risk, risk difference and number needed to treat from one 2x2 table — because an odds ratio of 6.00 can describe a relative risk of 1.50.
An educational tool. McNemar’s test detects a change in the marginal proportions and does not establish what caused it; regression to the mean and ordinary time trends produce the same table as a real intervention effect.
Published McNemar's test with the property that makes it surprising built in as a preset pair: the agreeing cells appear nowhere in the statistic. chi-square is (b − c)^2/(b + c), with no a and no d, so tables with n = 50, n = 1,425 and n = 25 produce byte-identical chi-square, p-value and exact p. The suite asserts identity rather than closeness.
That is correct rather than a defect — a subject who answered the same way twice says nothing about which direction change runs in — and it means the power depends on the discordant count rather than on the sample size. A study of ten thousand people in which eight changed has the power of a study of eight.
Reports the exact binomial p alongside the chi-square and highlights it below 25 discordant pairs, where the approximation is unreliable. The exact version is a two-sided binomial test on the discordant pairs at p = 0.5, which is all the test ever was.
Separates the odds ratio for change, b/c, from the marginal shift — two quantities that are routinely conflated, since the first is conditional on having changed at all and the second is what a reader usually means by the effect.
Add this calculator to your site
Responsive embed — and private: nothing your visitors type leaves their browser.