A scholarship awarded to students scoring 60 or above, and a later outcome. The true jump is 4.2. At the narrowest bandwidth the estimate is 0.246 with p = 0.927 — nothing at all. At the widest it is 5.088 with p below 0.0001. Same rows, same cutoff, six defensible bandwidth choices, and the estimate spans a factor of twenty-two while the verdict at 5% flips. The default setting returns 4.013, close to the truth, and there was no way to know in advance that it would be the one.
That verdict does not survive the bandwidth sweep. Across six defensible bandwidths on these same rows the estimate ranges from 0.246 to 5.454, and 3 of 6 clear the 5% level while the rest do not. Reporting any single one of them as the answer conceals a choice that determined it.
Jump at the cutoff
4.0134
95% CI 0.583 to 7.444
Across six bandwidths
0.25 to 5.45
3 of 6 significant
Placebo, left side
-0.5444
a fake cutoff where nothing should happen
Placebo, right side
-1.3121
a fake cutoff where nothing should happen
The same data at six bandwidths
Estimate, standard error, sample used and p-value at each bandwidth
Bandwidth
Rows used
Jump
Standard error
p
At 5%
4.88
31
0.2459
2.6490
0.9267
not significant
7.32
41
3.1650
2.2441
0.1668
not significant
9.76
57
3.3315
1.8078
0.0709
not significant
12.20
71
4.0134
1.7504
0.0250
significant
17.08
107
5.4536
1.4600
0.0003
significant
24.41
148
5.0883
1.2515
0.0001
significant
Narrow bandwidths use fewer points near the cutoff, which is less biased and noisier. Wide ones use more data further away, which is more precise and picks up curvature that has nothing to do with the cutoff. Neither end is correct and the trade has no data-free answer, which is exactly why the whole sweep belongs in a report rather than the one row that was chosen.
The two fitted lines
Intercept and slope of the local linear fit on each side of the cutoff
Side
Value at the cutoff
Slope
Rows
Below
52.1934
0.4949
33
At or above
56.2068
0.3400
38
Gap
4.0134
-0.1549
71
The jump is the gap between the two intercepts by construction, so those figures agree exactly on every run. The two slopes are allowed to differ, which is what keeps a bend near the cutoff from being read as a step.
The two placebo figures refit the same estimator at the median of each side, where no rule changes and no jump can exist. Values near zero are what a clean design produces; a large placebo jump says the outcome is simply lumpy in the running variable, and the estimate at the real cutoff is measuring that lumpiness.
A discontinuity design estimates the effect at the cutoff and nowhere else, and it assumes units could not precisely control which side they landed on. Where that fails, a policy change with a comparison group or an instrument may supply cleaner variation.
Six-bandwidth sweep Both fitted lines Placebo cutoffs Separate slopes each side A true-zero preset
What this tool shows
On the first preset the estimate runs from 0.246 to 5.454 across six defensible bandwidths — identical rows, identical cutoff — and the verdict at 5% flips from p = 0.927 to p = 0.0003. A twenty-two-fold range decided by one choice nobody outside the analysis ever sees. The second preset is sharper still: the true jump is exactly zero, five bandwidths correctly find nothing, and the sixth returns −2.76 at p = 0.029. Report that one and you have a discovery. Every run here prints the whole sweep, because a single estimate conceals the choice that produced it.
Local linear regression on each side, with separate slopes so curvature is not read as a step
The estimate at six bandwidths on every run, not one
A flag when the 5% verdict changes across that sweep
Placebo cutoffs at the median of each side, where no jump can exist
A preset where the true jump is zero and one bandwidth finds a significant effect
Both fitted lines, with the jump shown as the gap between their intercepts
Six-bandwidth sweep Both fitted lines Placebo cutoffs Separate slopes
One estimate hides the bandwidth that produced it.
Updated 13 September 2026 · Works in any browser, no installation
Regression discontinuity compares units just below a cutoff with units just above it, on the argument that whether someone scored 59 or 61 is close to arbitrary while the treatment they receive is not. Fitting a line on each side and reading the gap at the cutoff gives the effect. The difficulty is that “just below” and “just above” require a bandwidth, and the answer moves with it — on the first preset, by a factor of twenty-two.
At a glance
Formula shown
Within a window of width h either side of the cutoff, fit y = β₀ + β₁(x − c) + β₂·1[x ≥ c] + β₃·1[x ≥ c]·(x − c). The jump is β₂, which equals the gap between the two fitted values at the cutoff exactly. Narrow h means less bias and more noise; wide h means more precision and more contamination from curvature away from the cutoff. There is no data-free answer, so the estimate is reported at six values of h.
Scenario support
A scholarship awarded above a test score, a benefit that phases out at an income threshold, a class-size rule that triggers at an enrolment count, an election decided by a narrow margin, and any rule that switches on at a number.
Educational estimate
Planning support from the values you enter — not professional advice.
The choice that decides the answer
The first preset has a true jump of 4.2. Six bandwidths, all defensible, all applied to identical rows at an identical cutoff:
At the narrowest the estimate is 0.246 with p = 0.927 — no effect whatsoever. At 7.32 it is 3.165, p = 0.167. At 9.76, 3.332 and p = 0.071. At 12.20, 4.013 and p = 0.025. At 17.08, 5.454 and p = 0.0003. At the widest, 5.088 and p = 0.0001.
The estimate spans a factor of twenty-two and the verdict at 5% flips partway up. The default setting happens to return 4.013, very close to the truth, and there was no way to know in advance that it would be the one — the analyst choosing among these six sees only six numbers and no oracle.
The trade behind it is real rather than arbitrary. Narrow windows use points nearest the cutoff, where the comparison is most credible, and pay for it in noise: 31 rows instead of 148. Wide windows are more precise and quietly include units far from the cutoff, whose relationship to the outcome has nothing to do with the rule. Neither end is correct, and no amount of cleverness in choosing h makes the dependence go away. Printing the sweep does not solve the problem either. It makes the problem visible, which is the most an estimator can do about a choice that has to be made.
A significant jump where there is none
The second preset is the one worth sitting with. The true jump is exactly zero. The outcome is a straight line through the cutoff with ordinary noise on it, and the scholarship does nothing at all.
Five of the six bandwidths report that correctly, with p-values of 0.824, 0.937, 0.578, 0.627 and 0.294. The sixth returns −2.76 with p = 0.029, comfortably significant at the conventional level, on data containing no discontinuity.
That is not a bug and it is not unlikely. Six tests on the same data at a 5% level will produce a rejection reasonably often by chance, and the six are highly correlated so the arithmetic is not a simple multiple — but the practical point stands. An analyst who runs a sweep, sees one significant result, and reports that bandwidth has done nothing that looks like misconduct from outside. The write-up says a bandwidth was chosen and a jump was found.
The defence is not a better bandwidth rule. It is publishing the sweep, so a reader can see that five of six found nothing, and running placebo cutoffs where no jump can exist. Both are on every run here for that reason.
Cutoffs where nothing can happen
The two placebo figures refit the same estimator at the median of each side, well away from the real threshold. No rule changes there, so any jump found is the estimator responding to lumpiness in the outcome rather than to a treatment.
This catches a failure mode the bandwidth sweep cannot. If the outcome happens to be bumpy in the running variable — because of rounding, reporting thresholds, or a genuine nonlinearity — then a local linear fit will find steps in several places, and the one at the real cutoff is not distinguishable from the others on the strength of its size alone. A design with clean placebos and a jump at the cutoff is making a much stronger claim than one where every fake cutoff also jumps.
The third preset handles the related worry directly. There is no jump, but the relationship bends sharply near the cutoff, which is the classic way a smooth function gets mistaken for a step. Because each side gets its own slope, the local linear fits absorb the bend and all six bandwidths return nothing. Forcing a single straight line through both sides is what manufactures a discontinuity out of curvature, and it is why the two slopes are estimated separately here.
Whether units could choose their side
The design rests on the idea that landing at 59 rather than 61 is close to arbitrary. Where units can control that precisely, it is not, and the whole argument collapses — the people just above the cutoff differ from those just below in exactly the way that earns them the treatment.
The classic symptom is a pile-up of observations on the advantageous side of the threshold. A grade boundary where teachers can round up, an income threshold where earnings can be timed, a size rule a firm can stay under: in each case the density of the running variable jumps at the cutoff even though nothing about the underlying population does. Plotting that density is the standard check, and it is a check on the design rather than on the estimate.
A softer version of the same problem is a cutoff that triggers more than one thing. If scoring 60 brings a scholarship and also a different class, the jump measures both together, and no amount of bandwidth care separates them.
What the estimate applies to
The effect identified is the effect at the cutoff. Not the average effect of the scholarship, not the effect for a typical recipient — the effect for units sitting essentially at the threshold.
That is often the narrowest quantity in causal inference, and it is the price paid for its credibility. A student scoring 90 is not represented in this estimate at all, and there is no statistical route from what happens at 60 to what happens at 90 without assuming the effect is constant, which is exactly the sort of assumption the design was chosen to avoid.
It also means the estimate is not directly comparable to one from a policy change or a matched comparison, which describe different populations. Two credible studies of the same programme can disagree without either being wrong, because they are answering questions about different people.
Reporting it
Report the sweep, not the estimate. A table of bandwidths with the estimate, the rows used and the p-value at each lets a reader see immediately whether the finding is robust or was selected. On the second preset that table is the difference between an honest null result and a publication.
Report the placebo cutoffs and the density of the running variable around the threshold. Say how many rows fell inside the chosen window, because a discontinuity estimate resting on 31 observations is a different object from one resting on 148. And state plainly that the effect is local to the cutoff, since the number will otherwise be read as the effect of the programme.
Sources and methodology
References for the estimator, bandwidth selection and the manipulation test.
Method. The local linear estimator is checked against a global interacted regression at a bandwidth covering the whole sample, and the jump is verified to equal the gap between the two fitted intercepts on every run. Shifting the running variable and cutoff together is confirmed to leave the estimate unchanged. The bandwidth sweep and the placebo cutoffs are computed from the same fitted routine as the headline estimate, so the table and the tile cannot disagree. That engine is verified on every change against 556 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Difference-in-DifferencesEstimate a policy effect from treated and control groups before and after, with both naive comparisons, all four cell means and a pre-trend placebo test.
Instrumental VariableTwo-stage least squares with the first-stage F, the partial R-squared, the reduced form and the ordinary least squares estimate reported beside it.
Propensity ScoreFit a propensity model and read covariate balance before and after weighting, with common support and discrimination reported separately.
E-ValueHow strong an unmeasured confounder would have to be to explain an observed association away, for the point estimate and the interval limit.
Linear RegressionThe least-squares line with r and r² — and the regression of x on y beside it, because those are two different lines rather than one line rearranged.
Multiple RegressionFits several predictors with a VIF on every term, and names the configuration people misread: a model significant at p = 0.0103 where neither predictor reaches 0.05, at a VIF of only 7.11.
An educational tool. The estimate identifies the effect at the cutoff only and does not extend to units far from it. It assumes units could not precisely control which side of the threshold they landed on, which should be checked by plotting the density of the running variable. The bandwidth is a researcher choice that materially changes the estimate, and reporting a single bandwidth without the sweep conceals that dependence.