BOOK REVIEW · MEASUREMENT
Trustworthy Online Controlled Experiments
In short: the definitive book on online experimentation, written by the people who built testing culture at Microsoft, Google, and LinkedIn. Rigorous and practical, it covers the statistics, the pitfalls, and the organizational reality of trustworthy A/B tests.
What it covers.
The authors cover the full lifecycle of controlled experiments — design, statistics, common traps, and how to build an experimentation organization that trusts its results.
- Designing trustworthy experiments
- Statistics done right
- Common pitfalls and twyman’s law
- Guardrail and OEC metrics
- Experimentation at scale
- Building a testing culture
Who it’s for.
Anyone running A/B tests or building an experimentation program, plus measurement leads who need to trust test results. Technical but readable. Data scientists and product managers who own an experimentation platform will treat it as required reading.
Evergreen
- Trustworthiness over speed
- OEC and guardrail metrics
- The catalogue of pitfalls
- Experimentation culture
Read with a 2026 eye
- Heavier on web/product than media
- Assumes some statistics comfort
Key ideas worth stealing.
Their pitfalls catalogue (sample-ratio mismatch, peeking, Twyman’s law) is the most useful “what goes wrong” reference in measurement.
Defining an OEC (overall evaluation criterion) up front is the discipline that keeps experiments honest.
How it reads.
Textbook-rigorous but surprisingly readable, with war stories from real experiments at scale.
The RGM verdict.
The gold standard for experimentation — if your team runs tests, this prevents the expensive mistake of trusting broken ones. Pair with RGM’s statistical significance and incrementality guides. If your organization makes decisions from A/B tests, the cost of one untrustworthy result dwarfs the price of this book.
It is the reference you keep open while designing tests, not a one-time read.
Asked & answered.
Who wrote Trustworthy Online Controlled Experiments?
Ron Kohavi, Diane Tang, and Ya Xu — experimentation leaders from Microsoft, Google, and LinkedIn, often called the definitive authorities on A/B testing.
Is this book too technical?
It assumes some statistics comfort, but it is written for practitioners, not academics — the war stories and pitfalls make it readable.
How does it relate to incrementality testing?
Both rest on controlled experiments; this book is the rigorous foundation, while incrementality testing applies the same logic to media measurement.