Math calculator

I² Calculator

How much do studies disagree?

Measure the disagreement

The effects span 0.08 in total and τ is 0.0289 — the studies almost agree. I² is 89.32% anyway, because their standard errors are 0.01 and at that precision even a gap of 0.08 is far too large to be chance. High I² does not mean the studies disagree by much.

6 studies · Q = 46.8333 on 5 df

I² = 89.32%

95% CI 79.41% to 94.46%. τ = 0.0289 against an observed spread of 0.0800, and Cochran's Q gives p = 6.14375e-9. Admitting between-study variance multiplies the pooled standard error by 3.061.

89.32%

considerable

τ (between-study SD)

0.0289

τ² = 0.000837

Cochran's Q

46.8333

5 df, p = 6.1438e-9

H

3.0605

SE inflated 3.061×

I² is a percentage of the observed variation, not an amount. Here the studies span 0.0800 and the estimated real spread between them is τ = 0.0289. A high I² with a small τ means the studies are precise, not that they disagree by much.

Where Q comes from

Each study’s contribution to Cochran’s Q
StudyEffectDistance from poolQ contributionShare of Q
Study A0.5000-0.00830.69441.5%
Study B0.5300+0.02174.694410.0%
Study C0.4800-0.02838.027817.1%
Study D0.5500+0.041717.361137.1%
Study E0.4700-0.038314.694431.4%
Study F0.5200+0.01171.36112.9%
Total0.508346.8333100%

Each row is its inverse-variance weight times its squared distance from the fixed-effect pool. A precise study sitting a little way off contributes more than an imprecise study sitting far off, which is exactly why I² tracks precision rather than disagreement.

Interval on I², not just a point τ reported beside it Per-study Q contributions A ratio, never an amount

What this tool shows

Six studies whose effects span 0.08 in total give I² = 89.32%. Six studies spanning 2.00 give I² = 35.45%. Both presets are shipped below. Fifteen times as much real disagreement, less than half the I² — because I² is the share of the observed variation that is not sampling noise, not a measure of how far apart the studies are. τ is the one that answers that, and it is printed beside it.

  • I² with a confidence interval, which most heterogeneity calculators leave out
  • Cochran’s Q with its degrees of freedom and p-value
  • τ² and τ — the between-study variance on the scale of the outcome itself
  • H and the factor by which heterogeneity inflates the pooled standard error
  • Every study’s contribution to Q, so you can see which one is driving it
  • The truncation at zero made explicit, including when the raw ratio is negative
Interval on I² τ beside it Per-study Q Truncation shown

I² is a percentage of variation, never an amount.

Updated 13 September 2026 · Works in any browser, no installation

I² is the percentage of the variation across studies that is more than sampling error can explain. It is computed as (Q − df)/Q, where Q is the weighted sum of squared distances from the pooled effect. Because the weights are one over each study’s variance, a set of very precise studies reaches a high I² on differences that are practically meaningless, and a set of imprecise studies stays at a low I² on differences that are enormous. Read it with τ, never on its own.

At a glance

Formula shown
Q = Σwᵢ(yᵢ − θ̂)² with wᵢ = 1/SEᵢ² and θ̂ the fixed-effect pool; I² = max(0, (Q − df)/Q) × 100 on df = k − 1; H² = Q/df so that I² = (H² − 1)/H² exactly; and τ² = max(0, (Q − df)/C) with C = Σwᵢ − Σwᵢ²/Σwᵢ. The interval is built on ln H, where the sampling distribution is close enough to normal for a z interval, and mapped back through I² = (H² − 1)/H².
Scenario support
Deciding between fixed-effect and random-effects pooling, judging whether a set of trials belongs in one analysis, reporting heterogeneity for a systematic review, diagnosing which study is driving a high Q, and explaining to a reader why a very consistent literature can still show I² above 80%.
Educational estimate
Planning support from the values you enter — not professional advice.

High I² does not mean the studies disagree

This is the single most common misreading of I², and the two shipped presets make it impossible to miss.

Six precise studies: effects from 0.47 to 0.55, a total spread of 0.08. Their standard errors are 0.01. I² is 89.32% and τ is 0.0289.

Six imprecise studies: effects from −0.40 to 1.60, a total spread of 2.00. Their standard errors are 0.60. I² is 35.45% and τ is 0.4446.

The second set disagrees about fifteen times as much and scores less than half the I². Nothing has gone wrong. I² asks whether the variation exceeds what chance would produce, and at a standard error of 0.01 a gap of 0.08 is eight standard errors wide.

So “considerable heterogeneity, I² = 89%” can describe a literature that agrees to within 0.08. Quote τ alongside it and the sentence becomes honest.

The interval is usually far wider than the estimate suggests

I² is almost always reported as a bare number. It is an estimate from k data points, and at typical k it is a very loose one.

On the seven-trial preset I² is 58.00% with an interval from 2.88% to 81.83%. That range contains “might not be important” and “considerable” at the same time.

The width is driven by k, not by the data. On a set holding I² near 75%, the interval runs 0–91% at three studies, 40–90% at five, 50–89% at seven, 51–86% at ten and 61–84% at twenty.

Which makes the usual thresholds hard to use below about ten studies. An I² of 55% whose interval starts at 0% is not evidence of moderate heterogeneity.

Cochran’s Q has the mirror-image problem. It is underpowered at small k, so a non-significant Q is weak evidence of homogeneity — which is why the interval matters more than the test.

τ is the number with units

If only one heterogeneity statistic can be reported, τ is more useful than I² and it is reported far less often.

τ is the estimated standard deviation of the true effects across studies, on the same scale as the outcome. A τ of 0.0289 and a τ of 0.4446 mean something directly.

I² has no units and no scale. It cannot be compared between outcomes, and it cannot tell a reader whether the differences matter in practice.

τ is what drives the random-effects weights, and it is what the prediction interval in a meta-analysis is built from.

Its weakness is the opposite one: τ is badly estimated at small k, and the DerSimonian–Laird estimator used here is known to run low below about ten studies.

Which study is producing the heterogeneity

Q is a sum, so it can be broken apart, and the breakdown is usually more informative than the total.

Each row is that study’s weight times its squared distance from the pool. The shares add to 100% of Q by construction.

A single row holding most of Q means the heterogeneity is one study, not a pattern. That is a different finding from “the literature is inconsistent” and calls for a different response.

Precision decides who can contribute. A study with a tiny standard error that sits slightly off the pool contributes far more to Q than an imprecise study sitting far off, because the weight is one over the variance.

Which is the same mechanism as the paragraph above, seen per study: I² is driven by precision at least as much as by disagreement.

Why I² is so often exactly zero

An I² of exactly 0.00% is not a measurement. It is a truncation, and knowing that changes how to read it.

The raw ratio (Q − df)/Q goes negative whenever Q falls below its degrees of freedom, which happens routinely by chance.

On the fourth preset here, Q is 0.3679 on 5 df and the raw ratio is −1258.95%. It is reported as 0%, and τ² is likewise clamped to exactly zero.

So “I² = 0” means “no more variation than chance predicts”, not “the studies are identical”. Sometimes it means they agree slightly better than chance, which can itself be a signal worth examining.

When τ² is exactly zero the random-effects model reduces to the fixed-effect one, digit for digit — which is why the two models sometimes produce identical output.

The 25/50/75 thresholds are a rough guide

The familiar bands come from the Higgins and Thompson paper that introduced I², and that paper is explicit that they are tentative.

They are usually quoted as 25% low, 50% moderate, 75% high. The original wording is softer: “might not be important”, “moderate”, “substantial”, “considerable”, with overlapping ranges.

They carry no information about whether the differences matter. The 89% preset above sits in the top band on a spread of 0.08.

And they ignore the interval. A point estimate of 58% whose interval starts at 3% cannot be placed in a band at all.

Use them to start a conversation, not to end one — and never to choose between fixed and random effects after seeing the data.

Reporting heterogeneity

Four numbers, and the first two are the ones usually missing.

Give τ or τ². It is the only heterogeneity statistic on the scale of the outcome.

Give the interval on I². A bare percentage implies a precision that k studies cannot supply.

Give Q with its df and p. It is what I² and τ² are both derived from, and it lets a reader recompute either.

And give k. Every statistic on this page behaves differently below about ten studies, and a reader cannot judge any of them without it.

Sources and methodology

References for I², Q and the interval used here.

Method. Q is formed against the fixed-effect pooled estimate, which is what keeps I², τ² and the weights in a meta-analysis mutually consistent. I² is derived from Q rather than from H so that the published identity I² = (H² − 1)/H² can be asserted rather than assumed, and the suite checks it on 300 generated study sets. The confidence interval follows Higgins and Thompson on ln H, with the two standard-error branches either side of Q = k, and is withheld rather than faked when k is too small for the lower branch to exist. Both I² and τ² are truncated at zero, and the suite asserts that the truncation fires exactly when Q is at or below its degrees of freedom — never approximately. It also asserts that rescaling every effect and standard error by the same factor leaves Q and I² untouched, and that shifting every effect by a constant does the same, since heterogeneity is a property of the spread and not of the location. That engine is verified on every change against 134 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Meta-AnalysisPool study effects by inverse variance under both fixed-effect and random-effects models, with τ², a prediction interval, per-study weights and leave-one-out influence.
Egger's TestEgger's regression test for funnel-plot asymmetry, with the intercept, its interval, the plotted regression and I² beside it.
Chi-SquareGoodness of fit and tests of independence with every expected count and per-cell contribution shown — because the validity condition is about expected counts, not observed ones, and most calculators hide them.
Confidence IntervalIntervals for a mean or a proportion using t at every sample size and Wilson rather than the textbook Wald formula — with both methods shown, because Wald returns [0,0] at zero successes.
Standard ErrorStandard error of a mean or proportion, printed beside the standard deviation it gets confused with — the ratio is always √n, and at n = 50 that is a factor of seven.
Hedges' gBias-corrected standardised mean difference using the exact gamma correction, with Cohen's d beside it and a confidence interval.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool. I² is a ratio, not an amount: a literature that agrees to within 0.08 can score 89% if the studies are precise enough, and one that spans 2.00 can score 35% if they are not. Both I² and τ² are badly estimated below about ten studies, and both are truncated at zero, so an exact 0% is a clamp rather than a measurement.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (5 updates)

Published 13 September 2026

  1. Launched the I-squared calculator with a confidence interval on I-squared, which most heterogeneity tools omit entirely.
  2. Shipped two presets that separate precision from disagreement: effects spanning 0.08 give I-squared 89.32%, effects spanning 2.00 give 35.45%.
  3. Reported tau alongside I-squared throughout, since tau is the only heterogeneity statistic on the scale of the outcome.
  4. Added a per-study breakdown of Cochran Q so a single study driving the heterogeneity is visible.
  5. Made the truncation at zero explicit, including a preset where the raw ratio is -1258.95% and is reported as 0%.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.