Outlier
The point that stands apart. An outlier lies far from the rest of the data — a real rare event, an important signal, or an error to investigate before you act.
- Term
- Outlier
- Is
- A point far from the rest of the data
- May be
- A rare event, a signal, or an error
- Requires
- Investigation before action
Parts of speech & senses
- An outlier is an observation that lies an abnormal distance from the other values in a dataset, which may be a genuine rare event, an important signal, or an error worth investigating. "One outlier dragged the average way up."
What an outlier is
An outlier is a data point that sits far away from the rest of the values in a dataset — an observation so much larger, smaller, or otherwise unusual than its neighbors that it stands apart from the pattern. If most customers spend between ten and fifty dollars and one spends five thousand, that one is an outlier. Outliers can arise for very different reasons, and telling them apart is the whole art: some are genuine, rare-but-real events (a legitimately huge order, an extreme weather day); some are signals of something important (fraud, a system fault, an emerging trend); and some are simply errors — a mistyped figure, a broken sensor, a unit mix-up. The word describes the point's position relative to the rest of the data, not its cause.
Outliers matter because they can distort analysis out of all proportion to their number. A single extreme value can drag an average far from where the bulk of the data sits, inflate a measure of spread, or pull a model's fitted line toward itself, so conclusions drawn without noticing the outlier can be badly wrong. At the same time, outliers are sometimes the most important points in the dataset — the fraud you are trying to catch, the rare failure you must prevent, the exceptional customer worth understanding. So an outlier is neither automatically noise to discard nor automatically signal to chase; it is an anomaly that demands investigation before you decide which it is. Handling outliers thoughtlessly is one of the quiet ways analysis goes wrong.
Genuine outliers, errors, and detection
The first job with any outlier is to ask why it is there, because the right response depends entirely on the cause. A genuine outlier is a real, if rare, observation — it belongs in the data and may carry important information, so removing it would hide the truth. An erroneous outlier comes from a mistake — bad data entry, a faulty instrument, a processing bug — and it should be corrected or removed, because it does not reflect reality at all. The trouble is that both look the same on a chart: a point far from the rest. Distinguishing them takes investigation into the source of the data, not just its distance from the others. Deleting every outlier by reflex risks throwing away real signal; keeping every one risks polluting the analysis with errors.
Detecting outliers uses a mix of methods, from the simple to the statistical. Visual tools — scatter plots, box plots, histograms — reveal points that stand apart at a glance. Statistical rules flag points beyond a threshold, such as a number of standard deviations from the mean or outside the interquartile range. For complex, high-dimensional data, dedicated anomaly-detection algorithms find points that do not fit the learned pattern. Detection, though, only finds candidates; it does not tell you whether each is a real event or an error. That judgment still requires looking into where the point came from. The best practice is to detect systematically, then investigate each flagged point's cause before deciding to keep it, correct it, or set it aside.
Handling outliers well
Handling outliers well means investigating before acting. Detect them with appropriate visual and statistical methods, then trace each one to its cause: is it a genuine rare event, a meaningful signal, or an error? Correct clear mistakes, keep genuine observations (and consider whether they are the very thing worth studying), and be transparent about any point you exclude and why. Where a few extreme values would distort a summary, prefer measures that resist them — the median rather than the mean, robust models rather than fragile ones — instead of silently deleting the data. And treat some outliers as the point of the exercise: in fraud detection or quality control, the anomalies are exactly what you are hunting for, not noise to be swept away.
The failures are deleting outliers automatically to make the data look tidy (and losing real signal or hiding fraud in the process), keeping erroneous outliers that quietly corrupt every statistic, letting a single extreme value distort an average or a model without noticing, and never asking why an outlier exists. A related failure is treating outlier removal as a routine cleaning step rather than a judgment call that changes the results. The discipline is to detect outliers deliberately, investigate each one's cause, respond according to that cause rather than by reflex, use robust measures where extremes would mislead, and document what you did — because how you handle the anomalies often matters more to the conclusion than how you handle the ordinary points.
Synonyms & antonyms
Synonyms
Antonyms
Origin & history
An outlier — a data point lying far from the rest of a dataset, which may be a genuine rare event, a signal, or an error — must be investigated for cause before it is corrected, kept, or removed.
Etymology: source.
Usage trends
Search interest for this term over the last five years:
Common questions
- What is an outlier?
- A data point that lies far from the rest of the values in a dataset. It may be a genuine rare event, an important signal such as fraud, or simply an error — the word describes the point's position, not its cause.
- Should you always remove outliers?
- No. Deleting outliers by reflex can throw away real signal or hide the very thing you should study, like fraud. Investigate each one's cause first, then correct clear errors, keep genuine observations, and document any exclusion.
- How are outliers detected?
- With visual tools like scatter and box plots, statistical rules such as distance from the mean or the interquartile range, and anomaly-detection algorithms for complex data. Detection finds candidates, but their cause still needs investigation.
Resources & people to follow
- referenceRGM analysis — definitions, senses, and usage verified per term
Curated, non-competitor resources verified per term.
Related training
Disciplines
Areas of marketing where outlier is a core concern: