Statistical Power Analysis
Sizing the test before you run it. Power analysis works out how much data an experiment needs to catch an effect worth caring about — the antidote to underpowered tests.
- Term
- Statistical power analysis
- Is
- Sample-size planning for a test
- Power
- Chance of detecting a real effect
- Inputs
- Effect size, alpha, power
Parts of speech & senses
- Statistical power analysis is the process of calculating the sample size an experiment needs to detect an effect of a given size with a chosen probability. "Power analysis said we needed 20,000 sessions, not 2,000."
What power analysis is
Statistical power analysis is the planning step that decides how much data an experiment needs before it can trust its own answer. Statistical power is the probability that a test will detect an effect that is genuinely there — the chance of correctly rejecting a false null hypothesis, equal to one minus the Type II error rate. A power analysis turns that idea into a number by tying together four quantities: the effect size you want to be able to detect, the significance level (alpha) you will use, the power you want, commonly 80 or 90 percent, and the sample size. Fix any three and the fourth follows. Usually you fix the effect, alpha, and power, and solve for the sample size — the answer to 'how many visitors, sessions, or subjects do we need?' Done up front, it stops a test from being launched with no real chance of success.
Power analysis matters because the most common flaw in real-world testing is silent under-powering. A test with too small a sample cannot see a modest effect no matter how careful the rest of the analysis is; it returns 'no significant difference' and everyone treats a genuine win as a dud. That is a Type II error, and power analysis is how you prevent it. The single biggest driver of the required sample is the effect size: small effects need far more data to detect than large ones, which is why catching a half-point conversion lift can demand tens of thousands of visitors while a dramatic difference shows up in hundreds. Power analysis forces an honest conversation before any data is collected — how small an effect actually matters to us, and can we realistically gather enough traffic to catch it? If not, the test should be redesigned or not run.
Power analysis versus after-the-fact testing
There are two moments you can think about power, and only one of them is much use. An a priori power analysis happens before the experiment: you specify the effect you care about and compute the sample size needed to detect it, then collect exactly that much data. This is the valuable kind — it shapes the design and guarantees the test has a fair chance. A post-hoc power analysis happens after the fact, plugging the observed effect back in to report the power the test 'had.' This is largely circular and often misleading, because it just re-expresses the p-value you already have; a non-significant result will always look underpowered by this measure. The lesson is that power is a planning tool, not a consolation prize. Compute it before you run the test to size it correctly, not afterward to explain away a disappointing result.
Power analysis is the natural partner of significance testing, and each is incomplete without the other. A significance test, a t-test say, controls the false-positive rate through alpha — it protects you from seeing effects that are not there. Power controls the false-negative rate through beta — it protects you from missing effects that are. Run a significance test without having done a power analysis, and you have guarded one flank while leaving the other wide open. The two together define the whole risk profile of an experiment: alpha is how often you cry wolf, power is how often you catch the real wolf. Because tightening one can loosen the other for a fixed sample, the only clean way to hold both risks low is to compute the sample size that satisfies both — which is exactly what a power analysis delivers when you set alpha, power, and the effect size together.
Doing power analysis well
Doing power analysis well starts with an honest estimate of the smallest effect worth detecting — not the effect you hope for, but the smallest one that would change a decision. Anchor it in prior data, past experiments, or a clear business threshold rather than wishful thinking, because an inflated effect size produces a comfortably small sample that then misses the real, smaller effect. Set your alpha and target power deliberately; 0.05 and 80 percent are conventions, not laws, and higher-stakes tests deserve higher power. Feed in the baseline rate and the data's variability, since noisier metrics need more data. Then commit to the sample and duration the analysis returns, and resist the temptation to stop early. If the required sample is bigger than your traffic can supply, that is the analysis doing its job — telling you the test cannot answer the question as designed.
The failures are mostly ways of gaming the inputs or skipping the exercise entirely. Teams assume an optimistic effect size to justify a small, convenient sample, then wonder why real effects keep coming back null. They run tests with no power calculation at all and interpret every 'not significant' as 'no difference.' They lean on post-hoc power to rationalize a failed test instead of designing the next one properly. And they forget that variability and baseline rate drive the sample as much as the effect does, so they under-collect on noisy metrics. The discipline is straightforward: pick the smallest meaningful effect honestly, set alpha and power on purpose, compute the sample before collecting data, and treat an inconveniently large requirement as information, not an obstacle to argue around. Power analysis is what makes a test worth running before a single visitor is counted.
Synonyms & antonyms
Synonyms
Antonyms
Origin & history
The formal notion of statistical power comes from the Neyman–Pearson theory of hypothesis testing; systematic power analysis was popularized by psychologist Jacob Cohen from the 1960s.
Etymology: source.
Usage trends
Search interest for this term over the last five years:
Common questions
- What is statistical power analysis?
- A calculation that determines the sample size an experiment needs to detect an effect of a given size with a chosen probability (the power) at a chosen significance level, so a real difference is not missed.
- What is statistical power?
- The probability that a test detects an effect that is genuinely there — correctly rejecting a false null hypothesis. It equals one minus the Type II error rate, so higher power means fewer missed real effects.
- Why do power analysis before a test instead of after?
- Because power is a planning tool. An a priori analysis sizes the test to have a fair chance; a post-hoc analysis just re-expresses the result you already have and is largely circular, unable to rescue an underpowered study.
Resources & people to follow
- referenceRGM analysis — definitions, senses, and usage verified per term
Curated, non-competitor resources verified per term.
Related training
Disciplines
Areas of marketing where statistical power analysis is a core concern: