Math calculator

Difference-in-Differences Calculator

Two differences, subtracted.

Two differences

Employment across treated and control counties, four quarters, a policy arriving in the third. The true effect is −2.8. Looking only at the treated counties before and after gives +1.0779 — employment rose. Comparing treated against control after the policy gives +1.3889 — the treated counties are higher. Both readings are positive and the truth is negative, because employment was trending up everywhere and the treated counties started higher. The difference of the two differences returns −1.8943. Still 32% short of the truth, because this dataset also has a pre-trend the placebo test fails to find: it returns p = 0.8878.

112 observations · 2 pre-periods · four cells

Effect -1.8943 — the two naive readings say 1.0779 and 1.3889

The pre-trend placebo test returns p = 0.8878, which passes. That is weak evidence and not a licence: measured over 200 simulated draws, a pre-trend large enough to bias the estimate by 27% was caught by this test only 27.3% of the time. Passing it is compatible with a violation serious enough to produce the entire result.

Difference in differences

-1.8943

95% CI -3.346 to -0.443

Treated, before vs after

1.0779

opposite sign to the estimate

After, treated vs control

1.3889

opposite sign to the estimate

Pre-trend placebo

0.8878

passes — weak evidence only

The four cells the estimate is built from

Mean outcome in each group before and after, with the change and the difference of changes
GroupBeforeAfterChange
Treated103.9339105.01181.0779
Control100.6507103.62292.9721
Difference3.28321.3889-1.8943

The bottom-right cell is -1.89428571 and the regression interaction is -1.89428571. They agree exactly, and they must: four parameters fit four cell means with nothing left over, so the coefficient is the difference of the differences rather than an estimate of it.

What the placebo test can and cannot see

Pre-trend placebo test result and its measured power against violations of different sizes
QuantityValueReading
Pre-period slope gap0.1321per period, before the policy
Its standard error0.93212 pre-periods to estimate it from
Placebo p-value0.8878passes at 5%
Caught, bias 14%12.7%measured over 300 draws
Caught, bias 27%27.3%misses it three times in four
Caught, bias 61%78.0%only large violations are reliably found

Those three rows are simulated, not theoretical: 300 datasets at each violation size, counting how often the placebo test rejected. A pre-trend that shifts the estimate by more than a quarter goes undetected roughly three times in four. Passing this test is the absence of strong evidence against the assumption, which is a long way from evidence for it.

Parallel trends is an assumption about a counterfactual and is not testable, because it concerns what the treated group would have done after the policy. Where it is doubtful, a design with a different assumption may serve better — a cutoff rule, an instrument, or matching on measured covariates.

All four cells shown Both naive readings Placebo test with its power Pre-trend slope gap A true-zero preset

What this tool shows

On the third preset the true effect is exactly zero, and the method returns +2.7882 with p below 0.0001 and an interval that excludes zero. The treated group was simply on a steeper trend. And the pre-trend placebo test — the defence a referee would ask for — returns p = 0.0647 and passes. That is not a quirk of one dataset: simulating 300 draws, a pre-trend large enough to bias the estimate by 27% was caught by the placebo test only 27.3% of the time. Passing it is the absence of strong evidence against the assumption, not evidence for it.

  • The estimate with both naive comparisons printed beside it, because both are wrong by design
  • All four cell means, with the difference of the differences visible in the corner
  • A pre-trend placebo test, reported with its measured power rather than as a pass mark
  • A preset where the true effect is zero and every standard check is satisfied
  • A preset where both naive readings carry the opposite sign to the truth
  • The exact identity between the regression interaction and the raw cell means
All four cells Both naive readings Placebo test and its power Pre-trend slope gap

Parallel trends is an assumption about a counterfactual, so no test can confirm it.

Updated 13 September 2026 · Works in any browser, no installation

Difference-in-differences compares the change in the treated group with the change in a control group over the same period, and reports the difference between those two changes as the effect. Subtracting the control’s change removes anything that was happening to both groups; subtracting the before period removes any fixed gap between them. What survives is the effect, provided the two groups would have moved in parallel without the policy — which is the whole assumption, and it is about something that did not happen.

At a glance

Formula shown
The estimate is (ȳ_treated,after − ȳ_treated,before) − (ȳ_control,after − ȳ_control,before), equivalently the interaction coefficient in y = β₀ + β₁·treated + β₂·post + β₃·(treated × post). Those are the same number exactly, not approximately, because four parameters fit four cell means with nothing left over. The placebo test refits the same interaction on the pre-treatment periods only, where the coefficient should be zero if trends were parallel.
Scenario support
Evaluating a minimum wage change against neighbouring counties, a regulation that hit some firms and not others, a programme rolled out to some regions first, and any policy with a clear date, an affected group and a plausible comparison group.
Educational estimate
Planning support from the values you enter — not professional advice.

The test that passes anyway

The third preset was generated with a policy effect of exactly zero. Employment in the treated counties was on a steeper trend than the controls, and that trend simply continues through the policy date. Nothing happens when the policy arrives.

Difference-in-differences reports +2.7882. The p-value is below 0.0001 and the confidence interval runs from 1.5628 to 4.0136, nowhere near zero. Every conventional criterion for a finding is met, and there is nothing to find.

The pre-trend placebo test is the standard defence, and it is the reason this preset is worth printing. Refit the same interaction using only the periods before the policy, where the coefficient should be zero if the trends were parallel. It returns p = 0.0647. It passes at the conventional 5% level, with a couple of points to spare.

So the design is violated, the estimate is entirely artefact, and the diagnostic is clean. This is not a rare alignment. Simulating 300 datasets at each of several violation sizes and counting how often the placebo test rejected: a pre-trend biasing the estimate by 14% was caught 12.7% of the time; one biasing it by 27% was caught 27.3% of the time; only at a 61% bias did detection reach 78%. Under genuinely parallel trends the test rejects 6.75% of the time against its nominal 5%, so it is correctly sized — it simply has very little power, because two pre-periods is very little data with which to estimate a slope difference.

The first preset makes the same point on a dataset with a real effect. The truth is −2.8, the estimate is −1.8943 — 32% too small — and the placebo test returns p = 0.8878. It passes comfortably while missing a violation that cost a third of the answer.

Why both simpler comparisons fail

The design exists because the two obvious comparisons are both wrong, and the first preset shows them being wrong in the same direction at the same time. The true effect is −2.8.

Looking at the treated counties before and against after gives +1.0779. Employment rose. That reading attributes the entire common trend — everything that was happening to the whole economy — to the policy.

Comparing treated against control in the after period gives +1.3889. The treated counties are higher. That reading attributes the pre-existing gap between the two groups to the policy; those counties were already 3.3 points higher before anything happened.

Both are positive. The truth is negative. Subtracting one difference from the other removes the common trend and the fixed gap in a single step, and returns −1.8943. Still short, for reasons the placebo section covers, but at least pointing the right way — which neither simpler reading does. Both are printed on every run for exactly this reason.

Four cells, four parameters, no slack

The table of cell means has the estimate in its bottom-right corner, and that is not a coincidence of presentation. The saturated regression has four parameters — intercept, group, period, interaction — fitting four cell means. A model with as many parameters as cells reproduces those cells exactly, so the interaction coefficient is the difference of the differences, with no approximation anywhere.

Both are printed to eight figures on every run and they agree to 1.6e-14 across the verification suite, on balanced and unbalanced cells alike. It is worth seeing because it demystifies the regression: fitting an interaction sounds like a modelling choice with assumptions attached, and for the two-group two-period case it is arithmetic on four averages wearing different clothes.

What the regression form does buy is a standard error, covariates, and the ability to handle more than two periods or staggered adoption. Those extensions are genuinely different estimators with their own complications — in particular, the two-way fixed effects estimator can behave badly when units adopt at different times, because already-treated units end up serving as controls. This page handles the clean two-by-two case, which is where the intuition lives.

What parallel trends actually claims

The assumption is that the treated group would have moved parallel to the control group in the post-period had the policy not happened. It is a claim about a world that did not occur, and no amount of data from the world that did occur can confirm it.

The pre-trend test checks something adjacent and weaker: that the groups were parallel before. Groups can be parallel for years and diverge for reasons that have nothing to do with the policy — a recession hitting one industry, a separate programme, a shift in migration. They can also be non-parallel before and parallel after. The test is informative, and its informativeness is bounded in both directions.

One consequence is worth stating plainly: the assumption is not scale-invariant. If trends are parallel in levels they are generally not parallel in logs, and the reverse. Whether a policy effect is additive or multiplicative is a substantive judgement made before the estimator runs, and it changes what “parallel” means. Reporting both is more honest than picking whichever produces the cleaner placebo test.

Standard errors and the number of groups

The interval here treats observations as independent, which is right for the panel structure this page accepts and wrong for the most common real application. When many units are observed within a few treated and control clusters, outcomes inside a cluster are correlated and the conventional standard error is far too small.

The usual fix is clustering at the level of treatment assignment — the state, the county, the firm — and the usual catch is that it needs a reasonable number of clusters to work. With two treated units and two controls no adjustment saves the inference, because the effective sample size is four regardless of how many observations sit inside them. The same arithmetic is laid out on the design effect page, where the multiplier is computed directly.

This is the most frequent overstatement in applied difference-in-differences work, and it is entirely separate from the parallel-trends problem. An estimate can satisfy every trend diagnostic and still carry an interval that is several times too narrow.

Reporting it

Report all four cell means. A reader who can see them can reconstruct the estimate, spot which naive comparison would have misled, and judge whether the groups were remotely similar to begin with. Give the pre-period plot or the slope gap with its standard error, not just a placebo p-value, since a test with 27% power against a serious violation is not a result to lean on.

State the clustering level and the number of clusters. State whether the outcome is in levels or logs and why. If the placebo test passes, say so and say what its power was — describing a passing test as confirming parallel trends is the error the third preset is built to demonstrate, and it is one referees accept routinely.

Sources and methodology

References for the estimator, its standard errors and the limits of pre-trend testing.

Method. The interaction coefficient is checked against the raw difference of the four cell means on every run and agrees to 1.6e-14 across the suite. The placebo test’s size and power are measured by simulation rather than asserted: 200 draws under genuinely parallel trends give a 6.75% rejection rate against a nominal 5%, and 300 draws at each of four violation sizes give the detection rates quoted above. That engine is verified on every change against 556 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Regression DiscontinuityEstimate the jump at a cutoff by local linear regression, reported at six bandwidths with placebo cutoffs where no jump can exist.
Instrumental VariableTwo-stage least squares with the first-stage F, the partial R-squared, the reduced form and the ordinary least squares estimate reported beside it.
Propensity ScoreFit a propensity model and read covariate balance before and after weighting, with common support and discrimination reported separately.
E-ValueHow strong an unmeasured confounder would have to be to explain an observed association away, for the point estimate and the interval limit.
Design EffectThe design effect 1 + (m − 1)ρ with the effective sample size it implies, the ceiling of 1/ρ effective observations per cluster, and the marginal gain from each extra person.
Multiple RegressionFits several predictors with a VIF on every term, and names the configuration people misread: a model significant at p = 0.0103 where neither predictor reaches 0.05, at a VIF of only 7.11.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool. The parallel trends assumption concerns a counterfactual and cannot be verified from data; the pre-trend test provides weak evidence at best, with measured power near 27% against a violation biasing the estimate by a quarter. Intervals here assume independent observations and will be too narrow when treatment is assigned at a cluster level. Staggered adoption requires estimators beyond the two-by-two case shown here.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (5 updates)

Published 13 September 2026

  1. Published the estimator with both naive comparisons and all four cell means printed alongside.
  2. Shipped a preset where the true effect is zero, the estimate is +2.7882 at p below 0.0001, and the placebo test passes.
  3. Measured the pre-trend test rather than trusting it: 27.3% power against a violation biasing the estimate by a quarter.
  4. Showed the regression interaction equals the raw difference of four cell means exactly.
  5. Stated that intervals assume independent observations and are too narrow when treatment is assigned by cluster.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.