Math calculator

Equivalence Test Calculator

Testing that a difference is small, rather than failing to show it is large.

Is the difference small enough to ignore?

Ordinary p = 0.68 and TOST p = 0.022: no difference detected, and the difference shown to be small.

difference 0.5000 ± 1.2000, bounds ±3.0000

Equivalent — TOST p = 0.021828

The ordinary test detected no difference (p = 0.679150) — which on its own is not evidence that there is none.

TOST p

0.021828

the larger of two one-sided tests

Ordinary p (no difference)

0.679150

a different question

Lower one-sided p

0.002888

vs the lower bound

Upper one-sided p

0.021828

vs the upper bound

90% interval

-1.5206 to 2.5206

what TOST actually checks

Inside the bounds?

yes

Equivalence bounds

-3.0000 to 3.0000

Degrees of freedom

40

No difference detected, and the difference shown to be smaller than the bound. Both questions answered, and only the second is evidence of absence. A non-significant p on its own never is — absence of evidence is not evidence of absence, and TOST is the test that closes that gap by asking the question directly.
TOST is exactly the same as checking whether the 90% interval lies inside the bounds — and it does here. Note the level: a 90% interval at α = 0.05, not the 95% one you would report elsewhere. Two one-sided tests at 5% each correspond to a 1 − 2α interval, and using a 95% interval instead makes the procedure more conservative than it claims to be. The verification suite checks that the two agree on every one of 2,000 random configurations.
The bound is the hard part, and it is not a statistical question. ±3.0000is a claim about what size of difference would be unimportant in your subject, and it has to be set before the data — setting it afterwards is choosing the answer. Bioequivalence uses 80% to 125% on a log scale by regulatory convention; non-inferiority trials typically take a fraction of the effect the active comparator was shown to have. Where no convention exists, the bound is a judgement that should be stated and defended rather than derived.

What this tool shows

“No significant difference” is not evidence of no difference. The preset here has an ordinary p of 0.80 and still fails to establish equivalence at TOST p = 0.11 — the study detected nothing and demonstrated nothing. TOST asks the question directly, and it is exactly the (1 − 2α) interval check: a 90% interval at α = 0.05, not 95%.

  • Two one-sided tests against your equivalence bounds
  • The ordinary test of “no difference” beside it, since they answer different questions
  • All four outcomes named, including the one where both tests fail
  • The (1 − 2α) confidence interval, which TOST is exactly equivalent to
  • Both one-sided p-values, so the binding side is visible
  • The bound treated as the judgement it is, rather than as an input to guess
Equivalence, tested Four outcomes Interval identity Both p-values

p = 0.80, and equivalence still not shown.

Updated 12 September 2026 · Works in any browser, no installation

An equivalence test asks whether a difference is small enough to ignore. An ordinary test asks whether a difference can be detected, and failing to detect one is not the same as showing there is none — which is what a non-significant p-value is routinely read as saying.

At a glance

Formula shown
Two one-sided tests: H₀ that the difference is at or below the lower bound, and H₀ that it is at or above the upper one. Rejecting both concludes the difference lies strictly between them. The TOST p-value is the LARGER of the two one-sided p-values, because both must clear. Equivalently — exactly, not approximately — equivalence holds at level α if and only if the (1 − 2α) confidence interval lies entirely inside the bounds.
Scenario support
Bioequivalence of a generic against a reference drug, non-inferiority trials, showing a cheaper process matches the current one, demonstrating no meaningful difference between groups, and any question phrased as “are these the same” rather than “do these differ”.
Educational estimate
Planning support from the values you enter — not professional advice.

Why a non-significant result proves nothing

This is the most consequential misreading in applied statistics, and the tool is built to show it in two clicks.

A p-value above 0.05 means the data is consistent with no difference. It is also consistent with a large difference, if the study was imprecise enough. The test does not distinguish those cases and was never designed to.

The preset pair makes it concrete. The same observed difference of 0.5 gives an ordinary p of 0.68 at a standard error of 1.2, with equivalence established at TOST p = 0.022. At a standard error of 2.0 the ordinary p rises to 0.80 — a more reassuring-looking number — and equivalence fails at 0.11.

The second study is worse and its p-value is higher. That is the whole problem: a larger p can mean a smaller effect or a noisier measurement, and the number itself does not say which.

TOST asks the question directly. Not “can I rule out zero” but “can I rule out anything large”. Those are different hypotheses, and only the second supports a claim of no meaningful difference.

The confidence interval says the same thing, less formally. A narrow interval around zero is equivalence; a wide one is an uninformative study. TOST is that reading turned into a test with a stated error rate, which is why the two are exactly the same procedure.

It is the (1 − 2α) interval, not the 95% one

The identity is exact and the level is the part people get wrong.

Equivalence at α holds if and only if the (1 − 2α) interval lies inside the bounds. At α = 0.05 that is a 90% interval, not the 95% one reported everywhere else.

The reason is the two one-sided tests. Each spends α on one tail, so together they correspond to an interval leaving α in each tail — which is 1 − 2α in the middle.

Using a 95% interval instead is conservative, not wrong. It corresponds to α = 0.025 per side, so a study that clears it has demonstrated equivalence at a stricter level than claimed. It also fails to demonstrate it in cases where the correct procedure would have succeeded.

The tool prints the interval and the verdict together, and the verification suite checks they agree on every one of 2,000 random configurations, with both outcomes occurring so the check is not vacuous.

Which makes the interval the better thing to report. It contains the TOST result and the effect size and the precision, in one line, and a reader can apply their own bound to it.

Four outcomes, not two

Running both tests gives four possible conclusions, and only naming all four stops the ambiguous ones being written up as something cleaner.

Difference detected, not equivalent. A real effect, large enough to matter. The unambiguous case and the least common one.

No difference detected, equivalent. Also unambiguous: nothing found, and anything large ruled out. This is the finding people mean when they write “no difference”.

Difference detected AND equivalent. Not a contradiction. The effect is real and too small to care about, which happens routinely in large samples — with enough n any non-zero difference becomes detectable, which is why a significance test alone cannot answer a question about importance.

Neither. The study could not tell. This is the case that gets written up as “no significant difference” and read as equivalence, and it is the reason the page exists.

The tool names whichever case you are in, rather than reporting two p-values and leaving the combination to be interpreted. All four appear across the built-in presets.

The bound is a judgement, and it comes first

TOST takes one input that is not in the data, and everything depends on it.

The bound is the largest difference you would call unimportant. It is a claim about the subject matter — clinical relevance, economic materiality, engineering tolerance — and no amount of data supplies it.

It must be set before the analysis. Choosing it afterwards is choosing the answer, and a bound wide enough will declare anything equivalent. That is why regulators specify bounds in advance and why pre-registration matters as much here as anywhere.

Some fields have conventions. Bioequivalence uses 80% to 125% of the reference on a log scale, which is symmetric in log terms and is why the range looks lopsided. Non-inferiority trials typically take a fraction — often half — of the effect the active comparator was itself shown to have.

Where no convention exists, state and defend the choice. “We treated a difference under two points as clinically unimportant because the scale’s minimum detectable change is three” is an argument. A bound with no justification is an unexamined assumption doing all the work.

And the bound drives the sample size. A tighter bound needs a more precise study, in the same 1/d² way an effect size does — which the power calculator covers. Equivalence studies are typically larger than superiority ones for this reason.

More noise can lower the TOST p-value

The monotonicity most people assume holds only on one side, and the exception is worth understanding because it looks like a bug.

When the observed difference is inside the bounds, more noise raises the TOST p. That is the intuitive direction: a less precise study finds it harder to demonstrate equivalence. Verified on all 2,728 such cases in the suite.

When the difference is outside the bounds, more noise LOWERS it. On all 272 such cases. The extra noise is weakening the evidence against equivalence rather than for it.

A concrete pair: difference 4, bounds ±2. At a standard error of 0.5 the TOST p is 0.9999; at 2.0 it is 0.8383. Neither is anywhere near equivalent, which is why the reversal is harmless in practice.

The practical reading is that a noisy study can demonstrate nothing either way. It cannot show a difference and it cannot show equivalence, and the TOST p-value drifting toward the middle is that fact showing up in the arithmetic.

Which is another argument for the interval. A wide interval says “too imprecise” immediately, without needing anyone to reason about the direction a p-value moved.

Related procedures and where they differ

Several tests sit near this one and answer subtly different questions.

Non-inferiority uses one bound, not two. The question is whether the new treatment is not meaningfully worse, so only the lower bound is tested and the upper one is irrelevant. It is one of the two one-sided tests here, run alone.

A minimum-effect test is the mirror image. It asks whether the effect exceeds a threshold, rather than falls within one, and it is the right test when “statistically significant but trivial” is the failure mode you want to rule out.

Bayesian ROPE is the same idea in another framework. A region of practical equivalence around zero, with the posterior’s credible interval compared against it. Same bound, same judgement, different machinery.

Bayes factors can support a null directly. They quantify evidence for the null rather than failing to reject it, which is the other principled answer to the problem this page opens with — and they require a prior, which TOST does not.

Post-hoc power is not on this list. It is the procedure people reach for in exactly this situation and it carries no information beyond the p-value, as the power page shows. TOST is what post-hoc power is usually being asked to do.

And the simplest answer is often the interval. Reporting the effect estimate with its interval lets any reader apply their own bound, which is more useful than a single verdict computed against yours.

Sources and methodology

References for the procedure and for the reasoning behind it.

Method. The two one-sided tests use the t distribution on the degrees of freedom supplied, and the TOST p-value is the larger of the two, because both hypotheses must be rejected. The reported interval is the (1 − 2α) one rather than the (1 − α) one, which is the level the procedure actually corresponds to. The suite asserts that identity across 2,000 random configurations with zero disagreements and with both verdicts occurring, and it establishes the conditional monotonicity: more noise raises the TOST p-value on all 2,728 cases where the observed difference lies inside the bounds and lowers it on all 272 where it lies outside. That engine is verified on every change against 50 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Statistical PowerPower and sample size from the non-central t rather than a normal approximation, with the gap shown — plus a live demonstration that post-hoc power is a function of the p-value alone, and 0.500044 at p = 0.05 for every study ever run.
Confidence IntervalIntervals for a mean or a proportion using t at every sample size and Wilson rather than the textbook Wald formula — with both methods shown, because Wald returns [0,0] at zero successes.
p-valueA p-value from a t or z statistic, one- or two-tailed — with a panel that holds an effect fixed and grows the sample, so you can watch significance appear from nothing but n.
t-testOne-sample, two-sample and paired t-tests defaulting to Welch, with Student's pooled version printed beside it — and a warning when the two disagree on the verdict.
Effect SizeCohen d, Hedges g and the overlap between groups, with a sample-size control that moves the p-value while leaving the effect size fixed — the same d gives t = 1.29 at n=30 and 23.57 at n=10,000.
McNemar's TestChi-square and the exact binomial p for a paired 2×2 table, with the fact that makes the test surprising: the agreeing cells contribute nothing, so a study of 1,425 and one of 25 can give the identical result.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool and not medical or regulatory advice. The equivalence bound is a subject-matter judgement that must be set before the analysis; a bound chosen after seeing the data can declare almost anything equivalent.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (4 updates)

Published 12 September 2026

  1. Published an equivalence test that prints the ordinary 'is there a difference' p-value beside the TOST p-value, because the whole point is that both can fail. The preset pair: the same observed difference of 0.5 gives an ordinary p of 0.68 with equivalence established at 0.022, and with a wider standard error an ordinary p of 0.80 with equivalence NOT established at 0.11. The second study detected nothing and demonstrated nothing.
  2. Names all four possible outcomes rather than reporting two p-values and leaving the combination to be read, including the case where a difference is both detected and shown to be within the bounds — which is not a contradiction, since with enough sample any non-zero difference becomes detectable.
  3. States that TOST is EXACTLY the (1 − 2a) confidence interval check — a 90% interval at alpha = 0.05, not 95% — and verifies that identity across 2,000 random configurations with zero disagreements and both verdicts occurring.
  4. Establishes a conditional law the suite found rather than assumed: more noise raises the TOST p-value on all 2,728 cases where the observed difference lies inside the bounds and LOWERS it on all 272 where it lies outside, because there the extra noise weakens evidence against equivalence.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.