Sample Size
How long do we run it? is really how big a sample do we need? — and that's arithmetic, not a feeling.
- Term
- Sample Size
- Symbol
- n
- Calculated from
- Effect size, power, alpha, variance
- Not
- A fixed runtime or a round number
Forms & parts of speech
Definition in plain terms
Sample size is the number of observations (visitors, sessions, users) an experiment needs to detect an effect of a given size with acceptable power and confidence. It is the answer to 'how long do we run this?' translated into the question that actually has a calculable answer — and it's computed BEFORE the test from four inputs: the minimum effect worth detecting, the desired power (usually 80%), the significance level (usually 5%), and the baseline rate's variability.
The mechanics
The relationships are unforgiving: smaller effects need quadratically larger samples (halving the detectable effect roughly quadruples the required n), higher power and stricter significance both increase it, and higher baseline variance increases it. This is why testing tiny improvements on low-traffic pages can require impractical samples — sometimes the honest answer is 'this test isn't feasible.' The discipline that follows: fix the sample size in advance and run to it. Stopping early when results 'look significant' (peeking) inflates false positives badly — the fixed-horizon commitment is what keeps the math valid.
When it matters
Sample-size planning is the gate before every experiment — it converts a vague testing wish list into a feasible roadmap by revealing which tests your traffic can finish in a reasonable time. It matters most in low-volume contexts (B2B, expensive conversions) where adequate samples may take months, forcing a choice: test bigger swings, aggregate similar pages, or accept that some questions can't be answered by A/B testing and need other evidence. The number isn't bureaucracy — it's what separates a real experiment from theater.
Synonyms & antonyms
Synonyms
Antonyms
Origin & history
*Pieced together from statistical practice; the term has no recorded inventor. Sample-size determination grew from early-20th-century statistical theory (Fisher's experimental design, the Neyman-Pearson framework); A/B-testing platforms and writers like Evan Miller popularized accessible sample-size calculators that brought the discipline to digital marketing in the 2010s.
Etymology: source.
Usage trends
Search interest for this term over the last five years:
Common questions
- What is sample size in testing?
- The number of observations needed to detect a given effect with adequate power and confidence — calculated before the test.
- What determines required sample size?
- The minimum detectable effect, desired power, significance level, and the metric's baseline variance.
- Why not just pick a runtime?
- Validity depends on reaching the calculated sample and not stopping early — peeking when results look significant inflates false positives.
Related tools & calculators
Resources & people to follow
- bookTrustworthy Online Controlled Experiments — Kohavi, Tang & Xu
- referenceEvan Miller — sample-size and peeking essays
- referenceRGM analysis — sample size reveals which tests are even feasible
Curated, non-competitor resources verified per term.
Related training
- moduleCRO & experimentation
Disciplines
Areas of marketing where sample size is a core concern: