Building a culture of experimentation
Past a point, your results are limited not by tools or statistics but by culture — whether evidence actually beats opinion. This module covers the maturity model, how to win the executive sponsor who follows the test, making experimentation the default path, why celebrating losses unlocks the biggest wins, the skills a program needs, and scaling self-serve testing without letting velocity outrun trust.
What you will learn
Culture is the rate limit
Beyond a point, your experimentation results are limited not by tooling or statistics but by culture — whether the organization actually defers to evidence over opinion, tolerates the losses that learning requires, and gives teams the autonomy and cadence to test. The best testing tool in a HiPPO-driven culture produces nothing, because the highest-paid person overrides it. Culture is the rate limit on everything the prior five modules teach; this is where programs plateau or compound.
RGM tool · Diagnose your program with the CRO Maturity Score — it pinpoints your crawl/walk/run/fly stage and the weakest dimension to fix next.
The tell is what happens when a test contradicts a senior leader’s belief. In a mature culture, the data updates the belief; in an immature one, the data is explained away and the leader’s idea ships anyway. Every other capability — research, design, statistics — is downstream of that single cultural question: does evidence actually win here?
Getting numbers is easy; getting numbers you can trust is hard — and acting on them when they contradict your beliefs is harder still.
The maturity model
Experimentation maturity runs on a ladder: Crawl — occasional ad-hoc tests, no process, results ignored when inconvenient. Walk — a regular cadence, a real tool, some trustworthy results, but still HiPPO-vulnerable. Run — testing is the default way changes ship, statistics are respected, losses are accepted. Fly — experimentation is everywhere, many teams self-serve, every meaningful change is an experiment, and the org genuinely defers to evidence. Knowing your rung tells you the next realistic move — you can’t jump from Crawl to Fly.
Occasional tests, no real process, results ignored when they’re inconvenient. Experimentation is a side project nobody owns.
A regular testing rhythm, a proper tool, some trustworthy wins — but senior opinion still overrides results and only one team tests.
Meaningful changes ship via controlled tests, statistics are respected, losses are accepted as learning. The program reliably compounds.
Experimentation is org-wide, many teams run their own tests, nearly every change is an experiment, and leadership genuinely defers to evidence.
Claim: Mature experimentation organizations run thousands of controlled experiments per year and treat nearly every change as a test — the ‘Fly’ end of the maturity curve that companies like Microsoft, Google, and Booking.com operate at. Source: Kohavi, Tang & Xu — Trustworthy Online Controlled Experiments. Context: Maturity is a ladder, not a switch — diagnose your rung honestly and make the next realistic move, not a leap.
Org maturity models can feel abstract. There’s one question that locates you instantly: the last time a test result contradicted a senior leader’s strong opinion, what happened?
If the data updated the decision, you’re at Run or Fly. If it was explained away and the leader’s idea shipped anyway, you’re at Crawl or Walk no matter how good your tools are. That single anecdote predicts your program’s ceiling better than any tooling audit.
Fix the ‘evidence loses to seniority’ problem first; everything else is downstream of it.
Executive sponsorship
Experimentation culture is built top-down as much as bottom-up. A senior sponsor who publicly defers to test results — especially when a test kills their own idea — gives the program the air cover it needs to say no to pet projects and survive losing streaks. Without it, the first time a test contradicts a powerful stakeholder, the program loses, and everyone learns that evidence is optional. The single highest-leverage cultural move is converting one influential leader into a visible champion of ‘we test, and we follow the test.’
The way you win that sponsor is usually a high-stakes surprise: run a test on something leadership was sure about, and let a counterintuitive result — ideally one that saved money or avoided a bad launch — make the case for you. Nothing converts a HiPPO into a champion like watching their own confident assumption get disproven cheaply, before it shipped to everyone.
Process that makes testing the default
Culture is encoded in process. The defining shift is making the controlled experiment the default path for shipping meaningful changes — not an optional extra someone has to fight for. That means a standing research-to-backlog-to-test pipeline, a definition of which changes must be tested, a clear ship/kill decision protocol tied to the pre-registered metric, and a shared place where every result lives. When ‘did we test this?’ is a normal question in launch reviews, the culture has taken hold.
Culture change fails when testing is the harder option — an extra approval, a separate tool, a fight with the roadmap. People route around friction, so if testing is friction, people skip it.
So I engineer the defaults: a standing pipeline, templates for the hypothesis and analysis, pre-agreed which-changes-must-be-tested rules, and a one-click way to ship a change behind an experiment. The goal is that running the test is easier than not running it.
Behavior follows the path of least resistance — make the rigorous path the easy path and the culture changes itself.
Celebrating losses and learning
Since roughly two-thirds of ideas won’t win, a culture that punishes losing tests punishes experimentation itself — teams respond by testing only safe, trivial changes that can’t lose or teach. Mature organizations explicitly celebrate the learning, not just the wins: they share ‘what we learned’ from losses, reward bold well-designed tests regardless of outcome, and treat a disproven assumption as a save (a bad change caught before full rollout). Detaching ego and status from the win/loss outcome is what frees teams to test the big, uncertain ideas where the real wins hide.
Claim: With only about a third of experiments winning, treating losses as failures suppresses the bold tests that produce the biggest wins; high-performing programs reframe non-winners as learning and as prevented losses. Source: Kohavi et al. — experiment win-rate reality. Context: Reward the quality of the hypothesis and the test, not the outcome — or the team will only test safe, trivial changes.
Every leader says they celebrate learning from losses. Almost none do anything that makes it true, so teams correctly read the real incentive — don’t lose — and test timidly.
I make it concrete with a ritual: a ‘best loss of the quarter’ award for the well-designed test whose surprising negative result taught the most or prevented the costliest mistake. The team presents it; leadership applauds it; it goes in the deck.
Once a loss can win an award, the fear that drives safe, trivial testing evaporates — and the bold tests where real wins hide finally get proposed.
Skills and team
A program needs a blend of skills, not just a tool admin: research and analytics (to find what to test and read results honestly), statistical literacy (to avoid fooling yourselves), UX/design and engineering (to build trustworthy variants), and program management (to run the backlog and cadence). On small teams one person wears several hats; as you scale, a center-of-excellence model — a core team that sets standards, trains, and reviews while embedded teams execute — is the common way to keep quality high as volume grows.
Scaling from one team to the enterprise
Scaling experimentation is a quality-control problem: as more teams self-serve, the risk isn’t too few tests but too many bad ones — under-powered, peeked-at, SRM-broken — eroding trust in the whole program. Mature orgs scale with guardrails: shared tooling with automated validity checks (SRM, sample-size enforcement), an education program so every tester knows the unforgiving rules, a review or certification step for high-stakes tests, and a central team that sets standards. Velocity without trustworthiness is worse than a small, rigorous program.
- What’s the biggest barrier to experimentation culture?
- The HiPPO — the Highest-Paid Person’s Opinion overriding test results. If evidence loses to seniority, no tool or process matters. Executive sponsorship and a transparent, evidence-ordered backlog are the standard defenses.
- How do I get leadership to support experimentation?
- Convert one influential leader with a high-stakes surprise: test something they were sure about and let a counterintuitive, money-saving result make the case. Then make that leader the visible champion of ‘we test, and we follow the test.’
- Why celebrate losing tests?
- Because ~2/3 of ideas don’t win, and punishing losses makes teams test only safe, trivial changes. Celebrating the learning (and the bad changes prevented) frees teams to test the bold, uncertain ideas where the biggest wins are.
Where culture fails
Experimentation cultures fail when: the HiPPO overrides results, leadership claims to be data-driven but never lets evidence change its mind, losing tests are punished (so teams play it safe), testing is harder than not-testing (so people skip it), the org chases test volume without quality guardrails, and there’s no shared home for learnings so the same lessons get re-learned. All are organizational, and all cap the technical program below its potential.
Seniority beating evidence teaches everyone that testing is theater.
Leaders who never let data change their minds have a culture problem dressed as a tooling one.
Penalizing losses pushes teams to safe, trivial, unlosable (and unteachable) tests.
If the rigorous path has more friction, people route around it.
Scaling self-serve testing without quality control floods the org with untrustworthy results.
Your culture checklist
Culture is the rate limit — and it’s buildable. Tick what is genuinely true of your organization.
Prove it. Earn your passcode.
Ten questions, CASE method (Context · Analysis · Strategy · Execution). Pass at 90% to unlock this module’s completion passcode — retake as many times as you like.