Math calculator

Log-Rank Test Calculator

Two survival curves, compared.

Two curves, compared

Group 1 loses six of sixteen in the first six periods and then nothing. Group 2 loses nothing early and ten of sixteen after month 18. The running O−E peaks at +3.2652 and ends at −1.4717, so χ² = 0.5506 and p = 0.458065. The test sees almost nothing.

16 vs 16, 16 event times

χ²(1) = 0.550619, p = 0.458065

Group 1 had 6 events against 7.4717 expected; group 2 had 10 against 8.5283. The Peto one-step hazard ratio is 0.6879, with a 95% interval of 0.2561 to 1.8480 — which excludes 1 exactly when this test rejects, because log(HR)/se IS the log-rank z. The simpler (O₁/E₁)/(O₂/E₂) ratio gives 0.6848 and has no such guarantee. But the running observed-minus-expected CHANGES SIGN — it reaches 3.2652 at t = 6.000 before ending at -1.4717. The early and late differences are cancelling, and a single p-value cannot describe that.

χ²

0.55062

1 degree of freedom

p

0.458065

two-sided

Hazard ratio (Peto)

0.68788

95% CI 0.256 – 1.848

Peak |O − E|

3.2652

at t = 6.000, sign reverses

Observed minus expected, accumulating

-20201020timerunning O − E for group 1

This path crosses zero, which is the crossing-hazards signature. The final point — all the test uses — is much closer to zero than the path ever was.

Every event time with the at-risk counts, observed and expected events and the running total
TimeAt risk 1At risk 2Events 1Events 2Expected 1Running O − E
1.00001616100.5000+0.5000
2.00001516100.4839+1.0161
3.00001416100.4667+1.5495
4.00001316100.4483+2.1012
5.00001216100.4286+2.6726
6.00001116100.4074+3.2652
18.00001016010.3846+2.8806
19.00001015010.4000+2.4806
20.00001014010.4167+2.0639
21.00001013010.4348+1.6291
22.00001012010.4545+1.1746
23.00001011010.4762+0.6984
24.00001010010.5000+0.1984
25.0000109010.5263-0.3279
26.0000108010.5556-0.8835
27.0000107010.5882-1.4717

Group 1 ends at 62.50% survival and group 2 at 37.50%. Medians: not reached and 25.000.

Most powerful when the hazard ratio is constant Weights every event time equally Near-blind when the curves cross

What this tool shows

On the shipped crossing-hazards example the running observed-minus-expected reaches +3.2652 at t = 6 and finishes at −1.4717, so χ² = 0.5506 and p = 0.458065. One arm lost six of sixteen subjects in the first six periods while the other lost none; then the second lost ten of sixteen while the first lost none. The test uses only the final total, and the final total is nearly zero.

  • The log-rank χ² with its p-value, and the hazard ratio by the observed/expected method
  • Observed minus expected accumulated at EVERY event time, as a running total
  • The peak of that running total, and whether it changed sign
  • The full interval table: at-risk counts and events in each arm at each time
  • Kaplan-Meier survival and medians for both arms, for context
  • A crossing-hazards flag rather than a bare p-value
Two curves Running O − E Crossing flag Hazard ratio included

A non-significant log-rank test is not evidence that two curves agree.

Updated 13 September 2026 · Works in any browser, no installation

The log-rank test compares two survival curves by accumulating, at every time an event occurred, how many events one arm had against how many it would have had if the arms were identical. The total of those differences, scaled by its variance, is the statistic. It is the most powerful test available when one arm’s hazard is a constant multiple of the other’s — and it loses most of that power when the ratio changes over time.

At a glance

Formula shown
At each event time, E₁ = d·n₁/n where d is the total events and n₁, n are the numbers at risk, and V = d·(n₁/n)(n₂/n)(n−d)/(n−1). Then χ² = (ΣO₁ − ΣE₁)²/ΣV on 1 degree of freedom. Every event time contributes with equal weight, which is what makes the test optimal under proportional hazards and blind when early and late differences run in opposite directions.
Scenario support
Comparing treatment arms in a clinical trial, retention between two product cohorts, failure times for two component designs, time to relapse under two protocols, and any two time-to-event groups where some subjects have not had the event yet.
Educational estimate
Planning support from the values you enter — not professional advice.

Crossing hazards cancel, and the running total shows it

The test reduces the whole comparison to one number: the sum of observed minus expected across every event time. Anything that sums to zero is invisible to it, however dramatic.

The shipped preset is as dramatic as it gets. Group 1 loses six of sixteen in the first six periods and then nothing at all. Group 2 loses nothing for eighteen periods and then ten of sixteen.

The running total reaches +3.2652 at t = 6, then falls to −1.4717 by the end. The test uses only that last figure, which gives χ² = 0.5506 and p = 0.458065.

A bare p-value of 0.46 reads as “the arms are similar”, which is the opposite of what happened. The curves are about as different as two curves can be.

The tool plots the running total and flags the sign change, because that path is the diagnostic the statistic discards. A path that wanders far from zero and returns is the crossing-hazards signature; a path that stays on one side means the final total genuinely summarises the comparison.

It also happens without a sign change. The fourth preset peaks at 3.3702 and ends at 1.0188 — no crossing, but the total the test uses is a third of the largest separation, and p = 0.548516.

Equal weights are a choice, and there are alternatives

The log-rank test weights every event time the same. Two standard variants weight them differently, and each is more powerful for a different shape of difference.

Gehan-Breslow weights each time by the number at risk, so early differences count more — which suits a treatment whose benefit is immediate and wears off.

Tarone-Ware weights by the square root of the number at risk, a compromise between the two.

Fleming-Harrington lets you dial it, with parameters that emphasise early or late differences explicitly — which is the honest way to test a late-separation hypothesis, provided the choice is pre-specified.

Choosing the weighting after seeing the curves is the trap. Running all four and reporting the one that clears 0.05 is a selection with an error rate far above the nominal, and the log-rank test is the default precisely because it needs no such choice.

The hazard ratio it reports assumes what the test does not test

The tool prints a hazard ratio from the observed and expected counts, which needs no Cox model. It carries an assumption the log-rank test itself does not require.

The tool uses the Peto one-step form, exp((O₁−E₁)/V), because it is the one CONSISTENT with the test: log(HR) divided by its standard error is exactly the log-rank z, so the interval excludes 1 precisely when the test rejects. On the proportional-hazards preset that gives 3.1899 with a 95% interval of 1.0015 to 10.1602 against p = 0.049705 — both just clearing the line, as they must.

The simpler (O₁/E₁) ÷ (O₂/E₂) ratio gives 2.6554 on the same output, a 17% smaller number with no such guarantee. Both are printed, because a hazard ratio quoted beside a p-value that came from a different pivot is a quiet inconsistency.

Either way it is a summary over the whole follow-up, so quoting one presumes the ratio was roughly constant — the proportional-hazards assumption.

On the crossing preset it comes out 0.6879, a number that describes neither half of the study. The hazard ratio was enormously below 1 early and enormously above it late.

Which is why the crossing flag matters more than the hazard ratio does. When the path changes sign, no single ratio is a fair summary, and the honest report is the two curves plus survival at pre-specified timepoints.

The hazard ratio calculator goes into what the assumption costs and how to check it, including the test/interval disagreement that the observed-over-expected variance can produce.

Power depends on events, not on how many people you enrolled

This is the single most useful fact about designing a time-to-event study, and it follows directly from the formula: only event times contribute.

The variance is a sum over event times. A subject who never has the event contributes to the at-risk counts and to nothing else, so enrolling more people who will not have events buys almost no power.

Trials are therefore powered on the number of EVENTS, not the number of participants — roughly 4(zα+zβ)²/(ln HR)² events for a two-sided test, which is 508 events to detect a hazard ratio of 0.75 at 90% power and 380 at 80%.

Which is why event-driven trials stop at a pre-specified event count rather than at a calendar date, and why a rare outcome needs either a very long follow-up or a very large cohort.

And it is why a non-significant result with few events says little. Ten events across two arms cannot distinguish a hazard ratio of 1 from one of 2, whatever the enrolment was.

What it assumes beyond proportional hazards

Proportional hazards is about power rather than validity — the test keeps its error rate when they are not proportional, it just stops being able to see. Three other assumptions are load-bearing.

Censoring must be independent of prognosis, in both arms. Differential dropout — one arm losing sicker patients faster — biases the comparison, and nothing in the data reveals it. That is the same assumption Kaplan-Meier needs, applied twice.

Group membership must be fixed at time zero. Classifying subjects by something that happens later — responders against non-responders, say — builds the outcome into the grouping and produces a guaranteed difference. That is immortal time bias.

Subjects must be independent. Two eyes in one patient, or several devices from one batch, inflate the apparent event count and shrink the p-value.

And the chi-square approximation needs a reasonable number of events. Below about ten per arm an exact permutation version of the same statistic is the safer route.

Reporting a log-rank comparison

Four things, and the first is what turns a p-value into a finding.

Show the curves with the number at risk underneath. A reader cannot judge a late separation without knowing how many subjects were still there to separate.

Give the event counts per arm, not just the p-value. They determine the power, and “6 of 16 against 10 of 16” is a far more informative sentence than any statistic.

Say whether the hazards looked proportional. A crossing or converging pair makes the hazard ratio meaningless and the p-value uninformative, and it is visible on the plot.

And never report a non-significant log-rank test as showing the arms are similar. It shows the summed difference was small, which crossing curves also produce.

Sources and methodology

References for the log-rank test and its weighted variants.

Method. Expected events and variance are accumulated at every time an event occurred in either arm, using the hypergeometric form with the (n−d)/(n−1) correction so that tied event times are handled rather than approximated. The running observed-minus-expected is retained at every step, which is what lets the tool report the peak and the sign change rather than only the total — and it is the diagnostic the statistic itself throws away. The crossing flag is set when the running path is strictly positive at some point and strictly negative at another, which is a structural test rather than a threshold. Kaplan-Meier is computed for both arms from the same records so the survival context comes from one code path. Groups with fewer than two subjects, negative times, and a pair with no events at all return no result. That engine is verified on every change against 87 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Kaplan-MeierSurvival with censoring handled, and the naive count printed beside it: five events in twenty subjects give 25.0000% by the plain count and a 30.1202% cumulative incidence by Kaplan-Meier.
Hazard RatioBoth standard formulas from one log-rank output — Peto one-step 3.1899 and O/E 2.6554 — with a note on which of them is guaranteed to agree with the p-value beside it.
Life TableThe actuarial estimator with the effective-at-risk column shown and Kaplan-Meier beside it: ten-unit intervals cost 1.3753 percentage points, one-unit intervals cost exactly zero.
Relative RiskRisk ratio and odds ratio from one table with the divergence between them plotted: they agree to half a percent at a 1% baseline, and at an 80% baseline the odds ratio is exactly half the risk ratio.
Permutation TestEnumerates every split below 400,000 — two groups of ten is 184,756 of them — so the p-value is a ratio of two integers with no distribution assumed anywhere.
Chi-SquareGoodness of fit and tests of independence with every expected count and per-cell contribution shown — because the validity condition is about expected counts, not observed ones, and most calculators hide them.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool, not medical advice. The log-rank test sums observed minus expected across every event time, so two arms whose curves cross can produce a p-value near 0.5 while differing enormously — a non-significant result is not evidence that two curves agree. It also assumes censoring is unrelated to prognosis in both arms and that group membership was fixed at time zero.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (5 updates)

Published 13 September 2026

  1. Published a log-rank comparison that retains the running observed-minus-expected at every event time, not just the total the statistic uses.
  2. Built the page on the case the test is blind to: a running path that peaks at +3.2652 at t = 6 and finishes at −1.4717, giving χ² = 0.5506 and p = 0.458065 — while one arm lost six of sixteen subjects in the first six periods and the other lost ten of sixteen after month eighteen. The peak is more than twice the final total and opposite in sign.
  3. Flagged the sign change structurally rather than by threshold, and plotted the path, because that path is precisely the diagnostic the statistic discards.
  4. Switched the hazard ratio to the Peto one-step form, which is the one consistent with the test — log(HR)/se IS the log-rank z — and printed the simpler O/E ratio beside it so the 16.75% gap between them is visible.
  5. Verified that swapping the arms flips z, leaves χ² unchanged and inverts the hazard ratio, and that total expected equals total observed across both arms on every generated trial.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.