A bathroom scale reports 72.43 kilograms every morning. A sensor returns the same value to three decimal places. A model assigns 0.873 probability with impressive stability. The repetitions look scientific because they agree with one another. But agreement among readings is not the same as agreement with the quantity, event or outcome they are supposed to represent.
This is the quiet difference between precision and accuracy. Precision concerns how closely repeated measurements agree under stated conditions. Accuracy concerns closeness to the quantity value treated as true. A process can be precise but systematically offset. It can be roughly accurate on average but noisy from one reading to the next. It can also be neither.

Separate five ideas that everyday language compresses
Metrology uses a more careful vocabulary than ordinary product pages and dashboards. The International Vocabulary of Metrology defines measurement accuracy as closeness between a measured value and a true quantity value, while warning that accuracy is not itself a numerical quantity. It defines measurement precision through agreement among repeated indications or measured values under specified conditions.
Trueness concerns the agreement between the average of many measurements and a reference value. Repeatability is precision under a tightly controlled set of conditions, such as the same method, operator, system and location over a short interval. Reproducibility asks what happens when relevant conditions change. Resolution is the smallest change an instrument or display can distinguish or show. Uncertainty characterizes the dispersion of values that could reasonably be attributed to the measurand under a stated model.
These concepts interact, but they are not interchangeable. A display can have fine resolution and poor accuracy. A repeated process can have excellent short-term precision and drift across days. A corrected result can be close to the reference while still carrying material uncertainty about remaining effects.
| Question | Concept | What can go wrong |
|---|---|---|
| Do repeated readings agree? | Precision | All readings share the same systematic offset |
| Is the result close to a suitable reference? | Accuracy / trueness | The reference is wrong, unstable or irrelevant |
| Does the result hold when conditions change? | Reproducibility | Operator, device, environment or time reveals hidden variation |
| How small a change is displayed? | Resolution | Extra digits imply information the process does not contain |
| What range of values remains plausible? | Uncertainty | A point estimate is presented without its limitations |
Decimal places are a formatting decision before they are evidence
A number printed to six decimal places is not automatically known to six decimal places. Software can extend a calculation indefinitely. A sensor can convert an electrical signal into a long digital value. Neither act establishes that the input, calibration, sampling process or model supports every digit.
False precision appears when the representation is more exact than the knowledge behind it. Reporting an average commute as 31.7284 minutes may be mathematically reproducible from the recorded rows, yet the underlying start times, missing trips, route changes and sampling population may not justify hundredths of a second. The arithmetic is exact relative to the input file. The claim about commuters is not.
Rounding is therefore not merely cosmetic. It should communicate the scale the evidence can carry. Too few digits can hide a decision-relevant difference. Too many can manufacture authority. The correct choice depends on the measurement process and use, not a universal number of decimal places.
Consistency can preserve the same mistake
Suppose a scale is offset by two kilograms. If it returns nearly the same offset every time, it may be highly precise and poor in trueness. Averaging more readings reduces random scatter around the wrong center; it does not remove the offset. More data from the same biased process can make the wrong estimate look increasingly stable.
Systematic effects enter through calibration, zeroing, sampling, definitions, sensor placement, selection and the model that converts observation into a result. A survey can consistently reach the same unrepresentative population. A benchmark can reliably score the same narrow behavior. A dashboard can precisely count events whose definition changed without notice.
This is why the dataset construction audit and the benchmark audit begin before the final number. Precision inside a pipeline cannot repair a mismatch between the pipeline and the question.
Repeatability is conditional
“We got the same result” is incomplete until the conditions are named. Was the same instrument used? The same operator? The same software version? The same room, population, prompt, time of day and preprocessing? Agreement under unchanged conditions is valuable, but it may reveal less than agreement after the conditions that matter have varied.
NIST's measurement-process guidance separates repeatability, reproducibility and stability because different time scales and operating conditions expose different components of variation. A process can be quiet during one session and unstable across weeks. A result can reproduce inside one laboratory and fail when transferred to another environment.
For ordinary decisions, stage the check. Repeat the measurement without changing anything. Then change one plausible source of variation: location, device, operator, sample, day or method. If the conclusion moves, the condition belongs in the claim.
Accuracy requires a defensible reference
Accuracy is not available merely by comparing a reading with any convenient number. The reference must be suitable for the quantity and intended use, with its own traceability and uncertainty. In some domains the “true” value cannot be known exactly; laboratories use calibrated standards and documented chains of comparison. In social or behavioral settings, the target itself may be contestable.
A productivity score, sentiment label or “engagement quality” is not a physical measurand waiting to be read. Its definition embeds judgment. The system may measure its defined proxy consistently while the proxy remains an incomplete representation of the concept. That is a construct-validity problem, not a reason to abandon measurement.
The claim–evidence–inference worksheet helps expose that bridge. Write the observed quantity in one column and the conclusion in another. If they are not the same, name the inference between them.
A six-part measurement audit
- Name the measurand. State the quantity or property intended to be measured, not just the instrument output.
- Describe the method. Record device, version, sampling, operator, environment and transformation from observation to result.
- Check the reference. Identify the calibration or comparison value, its date, its applicability and its own uncertainty.
- Test repeatability and change conditions. Separate within-session agreement from variation across time, people, devices and contexts.
- Match digits to evidence. Report only the resolution and rounding the whole process supports.
- Carry uncertainty into the decision. Ask whether plausible variation would change the action. If it would, improve the measurement or choose a reversible step.
Do not collapse this audit into a single “accuracy percentage.” Measurement quality is multidimensional. A result fit for one decision may be inadequate for another. The reversibility test helps scale the evidence requirement to the cost of being wrong.
What this distinction can and cannot establish
Sourced fact: metrology distinguishes accuracy, precision, trueness, repeatability, reproducibility, resolution and uncertainty. Inference: public claims improve when they state which dimension was evaluated instead of calling a result “accurate.” Judgment: the acceptable uncertainty and rounding depend on the decision. Not claimed: a visual target metaphor captures every technical definition or replaces domain-specific measurement standards.
Precision is useful. It reveals stability under conditions someone can specify. The mistake is asking it to prove more: that the target was the right one, the reference was sound, the method traveled well or the displayed digits deserve belief.
Primary and authoritative sources
- BIPM/JCGM, International Vocabulary of Metrology: measurement accuracy, with linked entries for measurement precision, trueness, repeatability and reproducibility.
- JCGM 200:2012, International Vocabulary of Metrology, third edition, the complete vocabulary and definitions.
- NIST Technical Note 1297, for evaluating and expressing measurement uncertainty and for cautions about qualitative terminology.
- NIST/SEMATECH Engineering Statistics Handbook, Measurement Process Characterization, for repeatability, reproducibility, stability, calibration and uncertainty.
END OF TRANSMISSION 051
Keep the question. Test the model.
Choose the narrowest claim the evidence can carry, then leave room for revision.