Eight parts, three operators, two trials each. Gauge R&R is 0.73% by variance contribution and 8.55% by study variation, with 16 distinct categories. Both criteria pass comfortably, and the part-to-part variation is 99.27% of the total — which is what a measurement system is supposed to look like: almost everything you see is the parts, not the measuring.
By variance contribution the same system reads 0.73%, which is that figure squared: the two standard criteria are the same quantity on different scales and they can land in different bands. The system resolves 16 distinct categories, against a usual minimum of five. Reproducibility exceeds repeatability, so the operators differ more than the instrument does.
R&R % study variation
8.55%
acceptable
R&R % contribution
0.73%
the same figure squared
Distinct categories
16
five or more is usable
Interaction p
0.3208
no interaction detected
Where the variation lives
Variance components with their contributions and study variation shares
Source
Variance
SD
% contribution
% study variation
Repeatability (equipment)
0.003654
0.06045
0.31%
5.56%
Operator
0.004592
0.06776
0.39%
6.23%
Operator × part
0.000410
0.02025
0.03%
1.86%
Reproducibility
0.005002
0.07072
0.42%
6.50%
Gauge R&R
0.008656
0.09304
0.73%
8.55%
Part-to-part
1.174899
1.08393
99.27%
99.63%
The contribution column adds to 100% across repeatability, operator, interaction and part-to-part — variances add and standard deviations do not, which is why the last column does not add to anything. Reproducibility and gauge R&R are subtotals of the rows above them, not separate sources.
The analysis of variance behind it
Sums of squares for the two-way analysis of variance
Source
Sum of squares
Share
Part
49.37709
99.38%
Operator
0.15589
0.31%
Operator × part
0.06264
0.13%
Repeat error
0.08770
0.18%
Total
49.68332
100%
Those four sums must add to the total exactly, and the verification suite asserts it on sixty generated studies. The average-and-range method cannot produce this table at all, which is why it cannot separate an operator-by-part interaction from repeat error.
ANOVA, not average-and-range Interaction tested Both acceptance criteria Parts must span the range
What this tool shows
The operator-bias preset reads 17.26% by variance contribution — marginal — and 41.54% by study variation, which is a failure. They are the same number squared, and most reports quote one without saying which. Both are printed here, alongside the operator-by-part interaction that the average-and-range method cannot see at all: on the marginal preset it comes out at p = 0.0290.
Gauge R&R by the two-way ANOVA method, with the interaction term separated
Both acceptance criteria — percent contribution and percent study variation — side by side
Number of distinct categories, the resolution measure the percentages leave out
An F test on the operator-by-part interaction, which points at training rather than equipment
The full variance decomposition, with the contributions adding to exactly 100%
The sums of squares, which must decompose exactly and are shown so they can be checked
ANOVA method Interaction tested Both criteria Full decomposition
Percent study variation depends on the parts you chose. Pick them badly and the gauge looks bad.
Updated 13 September 2026 · Works in any browser, no installation
A gauge R&R study splits the variation you observe into the parts and the measuring. Repeatability is the same operator measuring the same part twice; reproducibility is different operators measuring the same part. Together they are the measurement system’s contribution, and everything left over is real differences between parts. The result decides whether a measurement can be trusted to sort parts against a specification — and it is a study about the instrument, run on parts chosen to test it.
At a glance
Formula shown
A two-way ANOVA with interaction gives MS_part, MS_operator, MS_interaction and MS_error. Then σ²_repeat = MS_error, σ²_operator = (MS_op − MS_int)/(p·r), σ²_interaction = (MS_int − MS_error)/r and σ²_part = (MS_part − MS_int)/(o·r). Reproducibility is operator plus interaction; gauge R&R is repeatability plus reproducibility. Percent contribution is σ²_GRR/σ²_total and percent study variation is σ_GRR/σ_total — which is the square root of the first, and the reason the two criteria have different thresholds.
Scenario support
Qualifying a measurement system before using it for process control, deciding whether a gauge can sort parts against a tolerance, diagnosing whether measurement problems come from the instrument or the operators, meeting an AIAG or IATF measurement-system-analysis requirement, and checking whether a new fixture actually improved anything.
Educational estimate
Planning support from the values you enter — not professional advice.
The two acceptance criteria are one squared
Every gauge study reports a percentage, and there are two different percentages with the same name. They routinely land in different bands.
Percent contribution is a ratio of variances. Percent study variation is a ratio of standard deviations. One is the square of the other.
On the operator-bias preset that is 17.26% and 41.54%. Against the usual 10/30 thresholds, the first is marginal and the second is a clear failure.
√0.1726 = 0.4154, so nothing is in disagreement except the scale. The thresholds differ too — 1%, 9% and above for contribution against 10%, 30% and above for study variation — and those are also squares of each other.
Quote both, or say which. A report saying “gauge R&R was 17%” is ambiguous by a factor that decides whether the gauge passes.
The term the average-and-range method cannot see
Gauge studies are still often run by the average-and-range method, which is arithmetically simpler and loses a real source of variation.
The operator-by-part interaction is operators disagreeing about some parts and not others. It is a different problem from operators reading consistently high or low.
Average-and-range folds it into repeat error, so the study blames the equipment for something the equipment did not do.
The ANOVA separates it and tests it. On the marginal preset the interaction has p = 0.0290 and accounts for 1.65% of the total variance.
A significant interaction points at training, fixturing or an ambiguous procedure — something about how the measurement is taken on particular parts, which no new instrument will fix.
Distinct categories, and why a percentage is not enough
The number of distinct categories answers a question the percentages do not: how many different values can this system actually tell apart?
It is 1.41 times the ratio of part standard deviation to gauge R&R standard deviation, rounded down, and the usual minimum is five.
On the unacceptable preset it is 1. The system cannot distinguish any part from any other — every reading is measurement noise.
On the acceptable preset it is 16, which is plenty for anything short of a capability study on a tight tolerance.
It moves with the parts, not just the gauge. A study run on eight nearly identical parts gives a low count from a perfectly good instrument, which is the trap in the next section.
The parts decide the answer as much as the gauge does
This is the most common way a gauge study goes wrong, and it produces a failing result from a gauge with nothing wrong with it.
Percent study variation compares the gauge against the parts in the study. Narrow the range of parts and the same instrument scores worse.
Parts pulled from one shift or one batch are usually far too similar. They understate the real process variation and inflate every percentage.
The parts must span the range the process actually produces. Ten parts covering the working range, not ten consecutive parts from the same run.
The alternative is to compare against the tolerance instead, which is percent tolerance rather than percent study variation — a different criterion that does not depend on which parts were sampled, and the right one when the gauge is used to accept or reject.
Repeatability or reproducibility — the fix differs
A failing gauge R&R is a number; the decomposition is what tells you what to do about it.
Repeatability dominating means the equipment. On the unacceptable preset it is 44.11% of the total variance against 10.47% for operators — a better instrument, a better fixture or more trials per part.
Reproducibility dominating means the people. On the operator-bias preset operators are 15.36% against repeatability’s 1.38% — the instrument is fine and the operators are not using it the same way.
Those are entirely different budgets. One buys hardware, the other writes a procedure and runs a training session.
And the interaction is a third case again, where the disagreement depends on which part is being measured.
What the study needs to be valid
The arithmetic assumes a particular design, and departures from it are not caught by the numbers.
Crossed and balanced: every operator measures every part the same number of times. A ragged design returns no result here rather than a plausible-looking one.
Randomised order, and blind where possible. An operator who remembers the previous reading produces a repeatability estimate that is far too good.
Ten parts, three operators, three trials is the usual standard, and smaller studies give unstable variance components rather than wrong ones.
Operators should be the ones who normally do the job. A study run by the two most careful people measures something other than the production measurement system.
And destructive tests cannot be done this way at all, because no part can be measured twice; those need a nested design with parts treated as batches.
Reporting a gauge study
Five items, and the first is the one that makes the percentage interpretable at all.
Say which percentage you are quoting. Contribution and study variation differ by a square root and have different thresholds.
Give the design: parts, operators, trials. The variance components are unstable on small studies and a reader needs to know how small.
Give the decomposition, not just the total. Repeatability against reproducibility decides what the fix is.
Give the number of distinct categories. It is the resolution question and the percentages do not answer it.
And say how the parts were chosen. A study on parts that do not span the process range fails a good gauge, and nothing in the output reveals it.
Method. The variance components come from a two-way analysis of variance with interaction, not from the average-and-range method, which cannot separate an operator-by-part interaction from repeat error. Negative variance estimates — which the method can produce when a mean square falls below the one below it — are truncated at zero rather than reported as negative, and that truncation is visible in the table as a zero contribution. Both acceptance criteria are printed together because they are the same quantity on different scales, and quoting one without saying which is ambiguous by a square root. The suite asserts that the four sums of squares decompose exactly to the total on sixty generated studies, that the four variance contributions add to exactly 100%, and that gauge R&R is exactly repeatability plus reproducibility with the total adding part variation — identities rather than tolerances. It also asserts that an unbalanced or ragged design, a study with one operator, and a study with a single trial per cell all return no result rather than a plausible-looking one. That engine is verified on every change against 103 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
CpkCp, Cpk, Pp and Ppk with the defect rates they predict and the rate actually observed — including the built-in case where Cp is 2.05, Cpk is 0.57 and a quarter of the sample is already out of spec.
Concordance CorrelationLin's concordance correlation with its exact decomposition into precision and accuracy, the scale and location shifts separated, and a bounded interval.
Intraclass CorrelationAll six ICC forms from one subject-by-rater matrix, with the rater means that drive them apart: one 8x3 matrix gives ICC(1,1) = 0.1277 and ICC(3,1) = 0.9852.
Bland-AltmanLimits of agreement with confidence intervals on the limits themselves, a proportional-bias test and the correlation printed beside them for contrast.
One-Way ANOVAThe full F table with eta and omega squared, plus every pairwise gap — because a significant F says something differs and never says which, and ten groups tested pairwise carry a 90% false-positive rate.
Standard DeviationSample and population standard deviation, plus variance, mean, median, quartiles, z-scores, outliers, and confidence intervals.
An educational tool. Percent study variation compares the measurement system against the parts in the study, so a study run on parts that do not span the process range will fail a perfectly good gauge, and nothing in the output reveals it. The method also requires a crossed, balanced design with repeated measurement of the same parts, so destructive tests need a nested design instead.