One value sits far from the rest. A sensor jumps, a customer takes ten times longer, a result lands outside the familiar range, or one study reports an effect its neighbors do not. The unusual point attracts two opposite impulses: delete it as bad data, or elevate it as the signal everyone else missed. Both moves are stories added to distance.
An outlier is an observation that appears unusually far from other observations under some comparison. It is not a diagnosis. The point may reflect a transcription error, a failed instrument, a legitimate member of a heavy-tailed distribution, a different subpopulation, a changed process, or a rare event with real consequences. The work is to distinguish those possibilities without letting the desired conclusion choose the rule.

Outlying is relative to a model
A value cannot be far from “the data” in the abstract. It is far under a particular scale, grouping, time window and model. A response time of eight seconds may be extreme among cached requests and ordinary among cold starts. A household income may look extreme in a local sample and expected in a deliberately mixed national sample. A measurement may be surprising on a linear scale and unsurprising after a justified transformation.
NIST defines an outlier as an observation lying an abnormal distance from other values in a random sample, then immediately notes that “abnormal” requires a judgment or consensus and that normal observations must first be characterized. That caution matters. A box plot, z-score or formal test can flag a point under assumptions. None can prove the physical or social cause of the point.
Before attaching a label, write the comparison class: same instrument, same operating state, same population, same unit, same time regime. If those conditions are mixed, the distant point may reveal a grouping mistake rather than a bad observation. This is one reason the dataset is not the world: the rows inherit the boundaries of collection.
Four explanations deserve separate tests
| Explanation | Evidence to inspect | Wrong shortcut |
|---|---|---|
| Recording or processing error | Source record, units, parsing, timestamps, duplicated transforms and manual entry | Deleting the value because it looks inconvenient |
| Measurement or procedure failure | Calibration, range, device status, protocol deviations and environmental conditions | Assuming the instrument was wrong without a fault trace |
| Legitimate tail or subgroup | Distribution shape, sampling design, covariates and comparable cases | Treating rarity as impossibility |
| Process change or new event | Time order, independent sensors, external records and repeat observations | Declaring discovery from one unexplained point |
The categories can overlap. A process change can push a valid sensor beyond its rated range. A subgroup can be underrepresented because collection failed more often for that group. A correct value can expose a coding assumption that silently truncated earlier cases. Preserve the chain from event to record before choosing a repair.
Deletion needs a reason other than extremity
NIST's outlier-detection guidance distinguishes a point known to be erroneous from one whose status cannot be determined. If the point is demonstrably miscoded, correction or exclusion can be justified. When it could be random variation or scientifically interesting, the guidance says not to simply delete it and suggests considering robust techniques.
The practical principle is simple: exclude because evidence identifies a failure of inclusion, not because the value changes the result. A temperature logged in Fahrenheit inside a Celsius column has a provenance-based correction. A response that violates a predeclared eligibility rule can be excluded under that rule. A result that makes the p-value cross a preferred threshold is not an exclusion criterion.
Write exclusion rules before inspecting outcomes when possible. If a rule must be developed after discovery, label it as post-hoc, show its effect and test it on fresh data. The same honesty applies outside research. A dashboard that removes “unusual” service failures may create a calm average by deleting the incidents users most need the system to explain.
Use methods that fit the question
A box plot is a useful visual convention, not a universal law of nature. A Grubbs test is designed for a single outlier in an approximately normal univariate dataset; NIST directs analysts who suspect multiple outliers toward other tests. Time series, spatial data, mixtures and heavy-tailed processes require different models because observations are not interchangeable draws from one bell curve.
Robust summaries can limit how much a small number of extreme values control an estimate. Medians, trimmed procedures or robust regression may help, depending on the estimand. But robustness does not erase the obligation to understand important extremes. If the question is average transaction time, a robust center may be useful. If the question is whether any transaction can exceed a safety limit, the extreme case may be the object of study.
Compare at least two honest analyses: the primary analysis under the stated rule and a sensitivity analysis under a plausible alternative. If the conclusion reverses when one ambiguous point moves, report that fragility. Do not hide it behind a single precise number. This follows the broader discipline that precision is not accuracy.
An outlier may be the process changing
Order matters. A single point in a shuffled table looks isolated. The same point at the beginning of a sustained shift may be an early warning. Plot observations in time, mark changes in software, suppliers, policy, environment and sampling, and inspect whether the point is independent of what follows.
Do not romanticize the first anomaly. Most surprising points do not announce a new scientific era. But systems fail when automatic filters are allowed to remove exactly the observations that contradict the expected range. NIST's handbook uses the historical ozone-hole episode as a warning about automated rejection: extremely low readings were screened as outliers before the phenomenon was recognized. The lesson is not that every extreme is profound. It is that a filter can embed yesterday's model so deeply that tomorrow's event becomes invisible.
A nine-step outlier audit
- Freeze the raw record. Preserve the original value, timestamp, unit, source and processing history.
- State the comparison. Name the population, state, window, scale and model under which the point appears unusual.
- Check transcription and transforms. Inspect units, parsing, joins, duplicates, rounding and conversions.
- Check the measurement chain. Review range, calibration, device status, procedure and environment without improvising hazardous testing.
- Plot context. Examine order, subgroups and the full distribution rather than one summary statistic.
- List competing causes. Include error, legitimate tail, subgroup, dependence and process change.
- Apply a justified rule. Use a predeclared or clearly documented criterion appropriate to the data assumptions.
- Run sensitivity analyses. Report material conclusions with the point included, corrected or excluded as defensible.
- Seek independent evidence. Use another instrument, source, sample or future observation that did not create the original story.
The audit scales with consequence. A personal habit log may need only provenance and a note. A medical, safety, legal, financial or scientific decision may require qualified review, validated procedures and formal statistical support. A checklist is not a license to override domain standards.
Separate fact, inference and speculation
Sourced fact: formal outlier tests have assumptions, and NIST cautions against simply deleting unexplained extreme observations. Inference: preserving provenance and showing sensitivity makes later judgment more auditable. Judgment: the cost of a false dismissal should influence how aggressively an anomaly is investigated. Speculation: an unexplained value may represent a new mechanism; until independent evidence arrives, that remains one model among several.
The unusual point does not deserve automatic belief. It deserves protection from automatic erasure. Hold it still long enough to learn what kind of question it is.
Primary sources and further reading
- NIST/SEMATECH Engineering Statistics Handbook: What are outliers in the data?, for the definition and need to characterize the comparison distribution.
- NIST: Detection of Outliers, for the distinction between known bad data, unexplained extremes and robust methods.
- NIST: Grubbs' Test for Outliers, for the test's normality and single-outlier scope.
- NIST: Histogram Interpretation—Symmetric with Outlier, including the ozone-data warning.
END OF TRANSMISSION 053
Keep the question. Test the model.
Choose the narrowest claim the evidence can carry, then leave room for revision.