Math calculator

Instrumental Variable Calculator

The first stage decides.

The first stage

Earnings against years of schooling, instrumented by distance to the nearest college. The true return is 1.5. Least squares gives 3.5628, badly inflated because able people get more schooling and earn more anyway. The instrument gives 2.9701 — a step toward the truth, and a small one — with a first-stage F of 5.266. Below the conventional threshold of 10, which matters: the interval is now 5.05 times wider than the least squares interval, and the estimate still sits most of the way to the biased answer. A weak instrument buys very little bias reduction at a large cost in precision.

120 rows · first-stage F 5.266 · below the threshold of 10

Instrumented estimate 2.9701 — least squares says 3.5628

That first-stage F is below 10, and the threshold is not a formality. Bucketing 4,000 simulated datasets by the F actually observed, the instrumented estimate landed further from the truth than plain least squares in 57.6% of samples below F = 5 and 17.6% between 5 and 10. Here the interval is 5.05 times wider than the least squares interval, so the correction is being bought at a real price and may not be a correction at all.

First-stage F

5.266

below 10 — treat with suspicion

Instrumented estimate

2.9701

95% CI 1.163 to 4.777

Least squares

3.5628

± 0.358 — the biased comparison

Interval width cost

5.05×

wider than least squares

Both stages, and what each one costs

First stage strength, reduced form, and the two estimates with their standard errors
QuantityValueStandard errorReading
First-stage F5.266weak instrument
Partial R² of the instrument0.0427share of the regressor it explains
Reduced form-0.80600.4872instrument against the outcome
Least squares3.56280.1826biased by the confounder
Two-stage least squares2.97010.9222p = 0.0017
Approximate surviving bias19.0%roughly 1/F, a rule of thumb only

The one-over-F figure is an expectation across repeated samples, not a promise about this one. On the first preset it says 19% of the bias survives, while the estimate actually sits about 71% of the way from the truth to the least squares answer. A rule of thumb describes the average behaviour of an estimator whose distribution has very heavy tails when the instrument is weak.

How often instrumenting makes things worse

Share of simulated samples in which the instrumented estimate was further from the truth than least squares, by observed first-stage F
First-stage FSamplesIV worse than least squares
Below 51,47557.6%
5 to 1058517.6%
10 to 205076.9%
20 to 505602.5%
Above 508730.0%

Four thousand simulated datasets with a known true effect, bucketed by the F the analyst would actually see. The highlighted row is the bucket this dataset falls in. Below F = 5 the “corrected” estimate is worse than the uncorrected one more often than not — and it arrives wearing a much wider interval, so the analyst pays twice.

With one instrument and no controls this estimate is algebraically the Wald ratio — the covariance of the instrument with the outcome divided by its covariance with the regressor. Both are computed in the verification suite and agree to 8.9e-16, so the number above does not depend on which route was taken to it.

A strong first stage says the instrument moves the regressor. It says nothing about whether the instrument affects the outcome through any other route, which is the assumption that cannot be tested at all. Where no credible instrument exists, a policy change or a cutoff rule may supply the variation instead.

First-stage F Least squares beside it Precision cost Measured failure rate Reduced form shown

What this tool shows

Bucketing 4,000 simulated datasets by the first-stage F an analyst actually sees, two-stage least squares landed further from the truth than plain least squares in 57.6% of samples below F = 5 — and in 0 of 873 samples above F = 50. The rule of ten is not folklore. Below it the correction is a coin toss that usually loses, and it arrives wearing a much wider interval, so the analyst pays twice. On the third preset an F of 0.010 produces an estimate of 31.0099 with an interval from −510.81 to 572.83 — 1,129 times wider than the interval it was meant to improve on.

  • Two-stage least squares with the ordinary least squares estimate printed beside it
  • The first-stage F and partial R-squared, reported before the estimate rather than after
  • How much wider instrumenting made the interval, as a multiple
  • The measured share of samples in which instrumenting made the answer worse, by F bucket
  • The reduced form, so the numerator of the ratio is visible
  • A preset with no usable instrument at all, which still returns a number
First-stage F Least squares beside it Precision cost Measured failure rate

A strong first stage is necessary and nowhere near sufficient.

Updated 13 September 2026 · Works in any browser, no installation

An instrument is a variable that moves the regressor you care about but affects the outcome only through it, so the part of the regressor the instrument explains is uncontaminated by the confounder. Two-stage least squares predicts the regressor from the instrument, then uses that prediction in place of the original. The method lives or dies on the first stage: if the instrument barely moves the regressor, the estimator is dividing by something close to zero and returns noise with the authority of a correction.

At a glance

Formula shown
Stage one regresses the endogenous variable x on the instrument z and any controls, giving fitted values x̂. Stage two regresses the outcome y on x̂. With one instrument and no controls this reduces exactly to the Wald ratio cov(y,z)/cov(x,z), which the suite confirms to 8.9e-16. The first-stage F tests whether z belongs in stage one at all; the conventional threshold is 10, and the share of bias surviving is roughly 1/F on average across samples.
Scenario support
Estimating the return to schooling using distance to a college, the effect of military service using a draft lottery, the effect of a treatment using the doctor a patient happened to see, and any setting where a genuinely arbitrary nudge pushes some units into treatment.
Educational estimate
Planning support from the values you enter — not professional advice.

When the cure is worse than the disease

The usual defence of a weak instrument is that the estimate is at least unbiased in large samples, merely imprecise. That is true asymptotically and badly misleading in practice, and the table on the tool measures why.

Four thousand datasets were simulated with a known true effect of 1.0 and a confounder that biases least squares upward to about 1.49. Each was bucketed by the first-stage F the analyst would actually observe, and the question asked of each sample was simply: was the instrumented estimate further from the truth than the uncorrected one?

Below F = 5, the answer was yes in 57.6% of 1,475 samples. The correction was worse than no correction more often than not. Between 5 and 10 it fell to 17.6%; between 10 and 20, to 6.9%; between 20 and 50, to 2.5%; and above 50 it was 0 of 873. The conventional threshold of 10 sits almost exactly where the failure rate stops being material, which is a reassuring thing to discover about a rule of thumb that is usually quoted without justification.

The reason is that a weak instrument pulls the estimator toward least squares rather than scattering it symmetrically. Two-stage least squares is a ratio, and as the denominator approaches zero the distribution develops tails so heavy that its mean does not exist. The estimator concentrates near the biased answer and occasionally throws out something enormous. The third preset is the second case: F = 0.010, an estimate of 31.0099, and an interval from −510.81 to 572.83.

What instrumenting costs when it works

Even a good instrument is not free, and the interval-width multiple on the tool prices it. The strong preset has a first-stage F of 275 and an instrument explaining 70% of the variation in schooling; least squares gives 2.3873, the instrument gives 1.7884 against a truth of 1.5, and the interval is 1.25 times wider. Most of the bias removed for a quarter more uncertainty is an excellent trade.

The weak preset prices the same trade badly. The F is 5.266, the estimate moves from 3.5628 only as far as 2.9701 against a truth of 1.5, and the interval is 5.05 times wider. A fifth of the distance travelled for five times the uncertainty.

This is worth stating because the instrumented estimate is often presented as simply better, with the precision cost relegated to a standard error nobody reads aloud. The multiple is on the tool so the trade is explicit: how much bias was removed, and how much precision was surrendered to remove it.

The assumption no diagnostic can reach

Everything above concerns instrument strength, which is measurable. The assumption that actually makes an instrument an instrument is the exclusion restriction — that it affects the outcome only through the regressor — and nothing in the data can check it.

Distance to a college is the classic example and the classic problem. It moves schooling, which is visible in the first stage. It may also proxy for urban labour markets, family income, or local industry, each of which affects earnings directly. A first-stage F of 275 says the instrument is strong. It says nothing whatever about whether it is valid, and a strong invalid instrument is more dangerous than a weak one because it produces a tight interval around the wrong number.

With more instruments than endogenous regressors an overidentification test becomes available, and it is worth running while being clear about what it does: it tests whether the instruments agree with each other, so it can detect that some are invalid but cannot confirm that any are valid. If they are all invalid in the same direction, they agree perfectly and the test passes.

The reduced form is printed on the tool for a related reason. It is the instrument regressed directly on the outcome, and it is the numerator of the ratio. If the reduced form is indistinguishable from zero, no amount of dividing by a small first stage produces a credible estimate — it only produces a large one.

Whose effect is being estimated

When the effect differs across units, an instrument does not recover the average effect in the population. It recovers the average among compliers — the units whose treatment status the instrument actually changed — and that group is defined by the instrument rather than by any question a reader is likely to have.

With distance to a college as the instrument, the compliers are people who would have gone to university if one were nearby and did not because one was not. That is a specific and fairly marginal group. Their return to schooling may be nothing like the return for people who were always going to attend, and an estimate of 1.7884 is an estimate about them.

This is not a defect, but it is frequently glossed over. Two valid instruments for the same regressor can give genuinely different answers without either being wrong, because they move different people. A disagreement between instruments is often read as evidence one is invalid when it may simply be evidence that the effect varies.

The ratio underneath

With a single instrument and no controls, two-stage least squares collapses to something very simple: the covariance of the instrument with the outcome, divided by its covariance with the regressor. The suite computes both routes and they agree to 8.9e-16.

Seeing the ratio form makes the weak-instrument problem obvious rather than technical. The denominator is the first stage. As it shrinks toward zero the ratio explodes, and small sampling errors in a near-zero denominator translate into enormous swings in the answer. The F statistic is testing exactly whether that denominator is distinguishable from zero, which is why it belongs above the estimate rather than in a footnote.

Reporting it

Report the first stage before the estimate: the F, the coefficient on the instrument, and the partial R-squared. Report the reduced form. Report the least squares estimate alongside, because the distance between the two is the correction being claimed and a reader cannot judge it otherwise.

Argue the exclusion restriction in words, since no statistic will do it for you. Say who the compliers are. And if the first-stage F is below 10, say so prominently rather than in a note — on the measurements above, that estimate is worse than the one it replaced in a majority of samples, and presenting it as the corrected figure inverts what the reader should take from it.

Sources and methodology

References for two-stage least squares, weak instruments and what an instrument identifies.

Method. Two-stage least squares is cross-checked against the Wald ratio cov(y,z)/cov(x,z) on simulated data and agrees to 8.9e-16, and the second-stage standard error is computed from residuals against the original regressor rather than the fitted one. The failure-rate table comes from 4,000 simulated datasets with a known true effect, bucketed by observed first-stage F. That engine is verified on every change against 556 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Difference-in-DifferencesEstimate a policy effect from treated and control groups before and after, with both naive comparisons, all four cell means and a pre-trend placebo test.
Regression DiscontinuityEstimate the jump at a cutoff by local linear regression, reported at six bandwidths with placebo cutoffs where no jump can exist.
Propensity ScoreFit a propensity model and read covariate balance before and after weighting, with common support and discrimination reported separately.
E-ValueHow strong an unmeasured confounder would have to be to explain an observed association away, for the point estimate and the interval limit.
Multiple RegressionFits several predictors with a VIF on every term, and names the configuration people misread: a model significant at p = 0.0103 where neither predictor reaches 0.05, at a VIF of only 7.11.
Linear RegressionThe least-squares line with r and r² — and the regression of x on y beside it, because those are two different lines rather than one line rearranged.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool. The exclusion restriction cannot be tested from data and must be argued substantively. Where effects vary across units, an instrument identifies the average among compliers rather than the population average. A first-stage F below 10 indicates a weak instrument, where the estimator is biased toward ordinary least squares and its sampling distribution has very heavy tails.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (5 updates)

Published 13 September 2026

  1. Published two-stage least squares with the first-stage F reported above the estimate rather than below it.
  2. Measured how often instrumenting makes the answer worse, bucketed by the F an analyst actually observes.
  3. Shipped a preset with no usable instrument that still returns a number, at an interval 1,129 times too wide.
  4. Priced the precision cost of instrumenting as a multiple of the least squares interval.
  5. Stated that the exclusion restriction cannot be tested and that effects are identified among compliers only.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.