Growth Marketing Glossary

Effect Size

ef·fect sizenoun

Not 'is it real' but 'how big' — the number that separates significant-and-trivial from significant-and-worth-shipping.

small effect - curves overlaplarge effect - clear separationnot 'is it real' - 'how big is it'the size of the difference, in standard units
Schematic — magnitude, separate from significance
Term
Effect Size
Asks
How big, not whether real
Forms
Cohen's d, relative lift, differences
Pairs with
P-values and confidence intervals

Forms & parts of speech

effect size · noun
The magnitude measure.
"p=0.001 and an effect size worth eleven dollars a month - real, and not worth the roadmap slot."

Definition in plain terms

Effect size measures how large a difference or relationship actually is — the question STATISTICAL SIGNIFICANCE never answers. A p-value reports whether a difference is distinguishable from noise; the effect size reports whether it is distinguishable from trivial. With enough sample, a 0.02-point conversion lift becomes solidly 'significant' while remaining commercially invisible — which is why every result this glossary's testing entries discuss deserves both numbers, and why the CHI-SQUARED entry's verdicts always travel with a lift attached.

The mechanics

The forms marketers meet: absolute differences (conversion up 0.55 points — the CONFIDENCE-INTERVAL entry's preferred companion), relative lifts (a 14% improvement on baseline), and standardized measures — Cohen's d expressing a difference in standard-deviation units (0.2 small, 0.5 medium, 0.8 large by Cohen's own deliberately rough benchmarks), correlation coefficients, and odds ratios in the modeling world. Standardization's job is comparability across metrics and studies; its trap is amnesia — a 'medium' d means nothing to a CFO until translated back into dollars, points, or customers, and the commercial translation IS the analysis's last mile. The operating roles: SAMPLE-SIZE planning runs on the minimum effect worth detecting (the test sized to find a 0.5-point lift, because smaller wouldn't change the decision — the chi-squared entry's pre-sizing ritual), meta-reading across experiments compares standardized effects where raw metrics differ, and the EXPERIMENT-readout discipline pairs every p-value with its effect and interval so 'significant' stops impersonating 'important.' The cultural failure it corrects has a name in the literature — the significance fallacy — and a budget cost in practice: roadmaps full of real-but-trivial wins shipped because the stars said so, while the practical-significance question (does this effect, at this size, justify this change?) went unasked.

When it matters

Effect size matters at every test readout — it is the half of the result that decisions actually run on — and upstream at test design, where the minimum-detectable-effect choice sets sample and duration. It matters most at scale, where significance comes cheap and triviality hides inside it. The discipline is the pairing rule (no p-value reported without its effect and interval), commercial translation as the final step, and the standing question pinned beside the dashboard's stars: big enough to act on?

Worked example. A marketplace's experimentation program celebrates a quarter of green dashboards - eleven 'significant winners' shipped - while the north-star metrics never move. The effect-size audit explains the paradox: at the platform's traffic, significance costs almost nothing, and the eleven wins average a 0.4% relative lift - real, trivial, and cumulatively invisible inside seasonal noise. The program rebuilds around magnitude: every test now pre-registers a minimum effect worth acting on (sized from the finance model - what lift pays for the engineering?), readouts pair p-values with effects and intervals translated to annual dollars, and the roadmap's bar moves from 'significant' to 'significant AND material.' Test volume drops by half; shipped impact triples. The stars still appear on the dashboards - they just stopped outranking the only question the business ever had: how big?
Failure modes to watch. Significance impersonating importance while triviality hides inside it; standardized effects quoted without commercial translation; tests sized with no minimum-effect decision, finding differences nobody would act on; readouts reporting stars without magnitudes; and roadmaps shipping real-but-trivial wins while material bets go untested.

Synonyms & antonyms

Synonyms

effect sizemagnitude (statistical)Cohen's d (one form)

Antonyms

p-value (the other half)statistical significance

Origin & history

Effect size's modern apparatus comes from Jacob Cohen's power-analysis tradition (his 1969 book fixed the d benchmarks and the discipline of sizing studies to detectable effects), and the replication-crisis decades made the pairing rule — effects with intervals, never bare p-values — the reporting standard marketing analytics inherited.

Etymology: source.

Usage trends

Search interest for this term over the last five years:

View interest-over-time on Google Trends →

Common questions

What is effect size?
A measure of how large a difference or relationship is — absolute differences, relative lifts, or standardized forms like Cohen's d — the magnitude question p-values never answer.
Why does effect size matter in testing?
With enough sample, trivial differences become 'significant' — effect sizes separate real-and-material from real-and-invisible, and set the minimum-detectable-effect that sizes tests.
How should effect sizes be reported?
Paired with p-values and confidence intervals, then translated commercially — dollars, points, customers — because 'medium effect' means nothing to a decision until it has units.

Related tools & calculators

Resources & people to follow

Curated, non-competitor resources verified per term.

Related training

Disciplines

Areas of marketing where effect size is a core concern:

Sources

  1. trendsGoogle Trends — "effect size"