Math calculator

Inverse Probability Weighting Calculator

Weights have a price.

What weights cost

Ninety trainees and a true programme effect of £4,000. Enrolment ran against prior earnings, so the people who took the programme were already poorer — and the raw comparison of outcomes returns −2.0319. The programme appears to have cost its participants two thousand pounds. Weighting each unit by the inverse of its enrolment probability gives +3.1688, the right sign and close to the right size. The interval still runs from −1.0282 to 7.3658, because weighting bought that correction with sample size: the effective size is 61.7 of 90.

90 units · 40 treated · 50 control

Weighted effect 3.1688 — the unweighted comparison says -2.0319

The heaviest weight is 10.402, carrying 5.93% of the total, and the effective sample size is 61.7 of 90 — a 31.4% efficiency loss. Note that the crude and weighted estimates have opposite signs: the raw comparison points the wrong way entirely.

Weighted effect

3.1688

95% CI -1.028 to 7.366

Unweighted

-2.0319

opposite sign

Effective sample size

61.7

of 90 — 31.4% lost

Heaviest single weight

10.40

row 73 — 5.93% of the total

What the weighting did, and what it cost

Estimates and weight diagnostics for the raw, weighted, stabilised and trimmed comparisons
ComparisonEstimateHeaviest weightEffective nNote
Unweighted-2.03191.00090confounded by design
Inverse probability3.168810.40261.7the headline estimate
Stabilised weights3.16885.77960.1differs by 7.11e-15
Trimmed at 1st/99th3.3620891 row dropped

Stabilising the weights shrinks the largest one from 10.402 to 5.779 and moves the estimate by 7.11e-15 — which is to say, not at all. The stabilising constants are the same within each arm, so they cancel in a difference of weighted means. They matter for a weighted regression; for this estimator they are cosmetic, and the page reports the measured difference rather than implying one.

Where the weight actually sits

Weight concentration and extreme propensity scores
DiagnosticValueReading
Effective sample size61.75the sample you actually have
Efficiency loss31.4%the price of the adjustment
Heaviest weight’s share5.93%no single row dominates
Scores outside 0.05–0.953few extreme units
Mean weight, treated2.0461.000 would mean no adjustment
Mean weight, control1.8721.000 would mean no adjustment

The effective sample size is Kish’s: the squared sum of the weights divided by the sum of their squares. It equals n exactly when every weight is identical and falls as they spread apart, which makes it the one number that reveals a study quietly resting on a handful of rows.

Weighting corrects for the covariates in the model and nothing else. Check that those covariates actually balance on the propensity score page before trusting a weighted estimate, and quantify what an unmeasured confounder would have to do with an E-value.

Effective sample size Efficiency loss Heaviest row named Stabilised and trimmed Raw estimate beside it

What this tool shows

On the third preset a single observation carries 66.85% of all the weight in the study, and the effective sample size falls to 2.2 out of 90. The estimate still prints, with a confidence interval that looks perfectly respectable. Nothing in the estimate reveals that two-thirds of it is one person. The first preset shows the other side: the unweighted comparison returns −2.0319 when the true effect is +4.0 — the wrong sign entirely — and weighting moves it to +3.1688. The correction is real, and it is bought with sample size.

  • The weighted estimate with the unweighted comparison printed beside it
  • Kish effective sample size and the efficiency loss, on every run
  • The heaviest single weight, which row it is, and what share of the total it carries
  • Stabilised weights, with the measured difference they make rather than an implied one
  • A trimmed estimate at the 1st and 99th percentile of the propensity score
  • A preset where weighting is nearly free, so the cost has something to be judged against
Effective sample size Efficiency loss Heaviest row named Stabilised and trimmed

A weighted estimate is only as large as its effective sample.

Updated 13 September 2026 · Works in any browser, no installation

Inverse probability weighting reweights each unit by one over its probability of receiving the treatment it actually received, so that the weighted sample looks like one in which treatment was assigned at random. A unit that was very unlikely to end up where it did stands in for many similar units that did not, so it gets a large weight. That is the mechanism, and it is also the failure mode: when a probability is near zero the weight is enormous and one observation can carry the study.

At a glance

Formula shown
Each unit is weighted by 1/e for the treated and 1/(1−e) for the controls, where e is the fitted propensity score, and the estimate is the difference of the two weighted means. The effective sample size is Kish’s: (Σw)²/Σw², which equals n exactly when the weights are identical and falls as they spread. Stabilised weights multiply by the marginal treatment probability — which cancels inside a difference of weighted means, so it changes the spread of the weights without changing this estimate.
Scenario support
Estimating a programme effect when enrolment depended on measured characteristics, adjusting for non-random attrition, standardising one population to another, and any comparison where treatment probability varies with covariates you have recorded.
Educational estimate
Planning support from the values you enter — not professional advice.

When one row takes over the study

The third preset is the case this page exists for. Enrolment was driven hard by both covariates, so some units had almost no chance of ending up in the group they are in. One of them receives a weight of 231.637. The total weight across all ninety units is such that this single row accounts for 66.85% of it — two-thirds of the answer is one person.

The Kish effective sample size makes that a number: 2.2 out of 90, a 97.5% efficiency loss. It is the squared sum of the weights over the sum of their squares, and it equals n exactly when every weight is identical. It falls as the weights spread, and a value of 2.2 says the weighted comparison has roughly the information content of a two-unit study.

The estimate is still printed: 8.3842, against a true effect of 4.0. More than twice the truth, with a confidence interval that reads perfectly normally. Nothing in the estimate, the interval or the p-value reveals what has happened. Fifty-one of the ninety units have propensity scores outside the 0.05 to 0.95 range, which is the underlying cause, and the effective sample size is the field that surfaces it.

The second preset gives the comparison that makes this legible. When enrolment was nearly random the heaviest weight is 2.413, the effective sample size is 89.4 of 90, and the efficiency loss is 0.7%. With proper overlap, weighting is close to free. Every warning here is about what happens when overlap fails, not about the method in general.

The correction is real

It would be easy to read the section above as an argument against weighting. It is not. The first preset has a true programme effect of +£4,000, and the raw comparison of outcomes returns −2.0319 — the programme appears to have cost its participants two thousand pounds. People who enrolled were already poorer, so their outcomes were lower whatever the programme did.

Weighting returns +3.1688. Right sign, close to the right size, from exactly the same rows. That is a large correction of a serious error, and the raw comparison would have supported the opposite conclusion with complete confidence.

The interval runs from −1.0282 to 7.3658 and covers zero, and that is the honest consequence of the effective sample size having fallen to 61.7. The estimate is no longer significant at 5%. A weighted analysis frequently trades a confidently wrong answer for a correctly uncertain one, and the uncertainty is the part that gets dropped when only the point estimate is reported.

Stabilised weights, measured rather than assumed

Stabilised weights multiply each raw weight by the marginal probability of the treatment that unit received, and the standard advice is that they behave better. On the weight distribution, they do: on the third preset the heaviest weight drops from 231.637 to 131.261, roughly halving.

On the estimate they do nothing whatsoever. The measured difference is 3.55e-14, which is rounding error. The reason is algebraic rather than empirical: the stabilising constant is the same for every unit within an arm, so it cancels between the numerator and denominator of each weighted mean. For a difference of two weighted means the two estimators are the same estimator.

This page prints that measured difference instead of quietly listing both numbers as though they were alternatives. Stabilisation earns its reputation in weighted regression and in marginal structural models, where the weights enter a fit rather than a ratio of sums, and the variance genuinely improves. Here it is cosmetic, and saying so is more useful than implying a choice that does not exist.

Trimming, and what it quietly changes

The usual remedy for extreme weights is trimming: drop units whose propensity scores fall outside some range, or cap the weights at a percentile. The table reports the estimate after trimming to the 1st and 99th percentile of the score so the effect of that choice is visible rather than assumed.

Trimming does reduce the variance, and it also changes the question. The units dropped are precisely those with no good counterpart, so the surviving estimate describes a narrower population — one that excludes the people least like the other group. That may be the right population to talk about. It is rarely the population the study set out to describe, and the change usually goes unmentioned.

An estimate that moves a lot under trimming is telling you that it rests on the units being trimmed. An estimate that barely moves is more reassuring, though it is not proof of anything: both the trimmed and untrimmed versions can be wrong in the same direction if the propensity model is wrong.

What weighting cannot do

Weighting removes confounding by the covariates in the propensity model. It does nothing about any other confounder, and it cannot tell from the data that one is missing. The weights will look perfectly reasonable, the balance table will be clean, and the estimate will be biased.

Two conditions have to hold beyond that. Every unit needs a genuine chance of either treatment — which is exactly what fails when scores reach 0 or 1 — and the propensity model has to be approximately right, since the weights inherit whatever the model got wrong. Neither is testable from the data alone, which is why the practical workflow runs through a balance table first and a sensitivity bound afterwards.

Where an assumption of this kind is unpalatable, the alternative is a design that does not need it: an instrument, a policy change with a comparison group, or a cutoff rule. Each replaces the no-unmeasured-confounding assumption with a different one, and none of them removes the need to assume something.

Reporting it

Report the effective sample size with the estimate. Not as an appendix figure — beside the interval, where a reader can see that a study of ninety had the information content of 2.2. Report the largest weight and its share of the total, and report the unweighted comparison too, because the distance between the two is the size of the correction being claimed.

State what was trimmed and how many rows it removed. If trimming changed the estimate materially, that is the finding, and reporting only the trimmed version presents a narrower population’s answer as the whole study’s. Balance diagnostics belong in the same table as the weights, because a weighted estimate whose covariates did not balance has no claim on anyone’s attention.

Sources and methodology

References for the weighting estimator and for the weight diagnostics.

Method. The weighted estimate is cross-checked in the suite against a hand-built weighted difference computed from the fitted scores directly, and the effective sample size is verified to equal n exactly under identical weights and to fall monotonically as they spread. The stabilisation result is measured, not asserted: the shift is below 1e-9 on every generated dataset while the largest weight strictly shrinks. That engine is verified on every change against 556 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Propensity ScoreFit a propensity model and read covariate balance before and after weighting, with common support and discrimination reported separately.
E-ValueHow strong an unmeasured confounder would have to be to explain an observed association away, for the point estimate and the interval limit.
Difference-in-DifferencesEstimate a policy effect from treated and control groups before and after, with both naive comparisons, all four cell means and a pre-trend placebo test.
Instrumental VariableTwo-stage least squares with the first-stage F, the partial R-squared, the reduced form and the ordinary least squares estimate reported beside it.
Design EffectThe design effect 1 + (m − 1)ρ with the effective sample size it implies, the ceiling of 1/ρ effective observations per cluster, and the marginal gain from each extra person.
Logistic RegressionLogistic regression with odds ratios converted to risk ratios at your own event rate, a likelihood ratio test, AUC, and separation reported rather than hidden.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool. Weighting removes confounding only by the covariates included in the propensity model and cannot detect an omitted one. It further requires that every unit had a genuine chance of either treatment, which fails when fitted scores approach 0 or 1, and the estimate inherits any misspecification in the propensity model.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (5 updates)

Published 13 September 2026

  1. Published inverse probability weighting with the Kish effective sample size on every run.
  2. Named the heaviest weight, its row, and its share of the total weight in the sample.
  3. Shipped a preset where one observation carries 66.85% of the study and the effective size falls to 2.2 of 90.
  4. Reported the measured shift from stabilised weights rather than implying they change the estimate.
  5. Added a good-overlap preset where weighting costs 0.7% of the sample, so the failures have a benchmark.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.