RGM-204 · CRO & Experimentation · Module 6 of 6

Building a culture of experimentation

Past a point, your results are limited not by tools or statistics but by culture — whether evidence actually beats opinion. This module covers the maturity model, how to win the executive sponsor who follows the test, making experimentation the default path, why celebrating losses unlocks the biggest wins, the skills a program needs, and scaling self-serve testing without letting velocity outrun trust.

What you will learn9 sections

Culture is the rate limit

Beyond a point, your experimentation results are limited not by tooling or statistics but by culture — whether the organization actually defers to evidence over opinion, tolerates the losses that learning requires, and gives teams the autonomy and cadence to test. The best testing tool in a HiPPO-driven culture produces nothing, because the highest-paid person overrides it. Culture is the rate limit on everything the prior five modules teach; this is where programs plateau or compound.

RGM tool · Diagnose your program with the CRO Maturity Score — it pinpoints your crawl/walk/run/fly stage and the weakest dimension to fix next.

The tell is what happens when a test contradicts a senior leader’s belief. In a mature culture, the data updates the belief; in an immature one, the data is explained away and the leader’s idea ships anyway. Every other capability — research, design, statistics — is downstream of that single cultural question: does evidence actually win here?

Getting numbers is easy; getting numbers you can trust is hard — and acting on them when they contradict your beliefs is harder still.
Adapted from Ronny Kohavi, Trustworthy Online Controlled Experimentsexperimentguide.com

The maturity model

Experimentation maturity runs on a ladder: Crawl — occasional ad-hoc tests, no process, results ignored when inconvenient. Walk — a regular cadence, a real tool, some trustworthy results, but still HiPPO-vulnerable. Run — testing is the default way changes ship, statistics are respected, losses are accepted. Fly — experimentation is everywhere, many teams self-serve, every meaningful change is an experiment, and the org genuinely defers to evidence. Knowing your rung tells you the next realistic move — you can’t jump from Crawl to Fly.

Crawl — ad hoc

Occasional tests, no real process, results ignored when they’re inconvenient. Experimentation is a side project nobody owns.

THE MOVE · Next move: pick one owner, one funnel area, and establish a basic cadence and trustworthy measurement — prove value once.
Walk — cadence forming

A regular testing rhythm, a proper tool, some trustworthy wins — but senior opinion still overrides results and only one team tests.

THE MOVE · Next move: secure executive sponsorship and a transparent backlog so evidence, not the HiPPO, orders the queue.
Run — testing is default

Meaningful changes ship via controlled tests, statistics are respected, losses are accepted as learning. The program reliably compounds.

THE MOVE · Next move: enable more teams to self-serve, invest in shared tooling/governance, and start celebrating losses publicly.
Fly — embedded everywhere

Experimentation is org-wide, many teams run their own tests, nearly every change is an experiment, and leadership genuinely defers to evidence.

THE MOVE · Next move: guard quality at scale (review, SRM automation, education) so velocity doesn’t outrun trustworthiness.

Claim: Mature experimentation organizations run thousands of controlled experiments per year and treat nearly every change as a test — the ‘Fly’ end of the maturity curve that companies like Microsoft, Google, and Booking.com operate at. Source: Kohavi, Tang & Xu — Trustworthy Online Controlled Experiments. Context: Maturity is a ladder, not a switch — diagnose your rung honestly and make the next realistic move, not a leap.

RGM EXPERT TRICK
Diagnose your rung by what happens when a test contradicts the boss

Org maturity models can feel abstract. There’s one question that locates you instantly: the last time a test result contradicted a senior leader’s strong opinion, what happened?

If the data updated the decision, you’re at Run or Fly. If it was explained away and the leader’s idea shipped anyway, you’re at Crawl or Walk no matter how good your tools are. That single anecdote predicts your program’s ceiling better than any tooling audit.

Fix the ‘evidence loses to seniority’ problem first; everything else is downstream of it.

WHY IT’S RARE · Leaders love to claim they’re data-driven. The honest test is whether evidence has ever changed their mind in public — that’s the real maturity signal.

Executive sponsorship

Experimentation culture is built top-down as much as bottom-up. A senior sponsor who publicly defers to test results — especially when a test kills their own idea — gives the program the air cover it needs to say no to pet projects and survive losing streaks. Without it, the first time a test contradicts a powerful stakeholder, the program loses, and everyone learns that evidence is optional. The single highest-leverage cultural move is converting one influential leader into a visible champion of ‘we test, and we follow the test.’

The way you win that sponsor is usually a high-stakes surprise: run a test on something leadership was sure about, and let a counterintuitive result — ideally one that saved money or avoided a bad launch — make the case for you. Nothing converts a HiPPO into a champion like watching their own confident assumption get disproven cheaply, before it shipped to everyone.

Process that makes testing the default

Culture is encoded in process. The defining shift is making the controlled experiment the default path for shipping meaningful changes — not an optional extra someone has to fight for. That means a standing research-to-backlog-to-test pipeline, a definition of which changes must be tested, a clear ship/kill decision protocol tied to the pre-registered metric, and a shared place where every result lives. When ‘did we test this?’ is a normal question in launch reviews, the culture has taken hold.

RGM EXPERT TRICK
Make ‘ship it as an experiment’ the path of least resistance

Culture change fails when testing is the harder option — an extra approval, a separate tool, a fight with the roadmap. People route around friction, so if testing is friction, people skip it.

So I engineer the defaults: a standing pipeline, templates for the hypothesis and analysis, pre-agreed which-changes-must-be-tested rules, and a one-click way to ship a change behind an experiment. The goal is that running the test is easier than not running it.

Behavior follows the path of least resistance — make the rigorous path the easy path and the culture changes itself.

WHY IT’S RARE · Most culture initiatives exhort people to test more. Removing the friction so testing is the default action is what actually changes behavior at scale.

Celebrating losses and learning

Since roughly two-thirds of ideas won’t win, a culture that punishes losing tests punishes experimentation itself — teams respond by testing only safe, trivial changes that can’t lose or teach. Mature organizations explicitly celebrate the learning, not just the wins: they share ‘what we learned’ from losses, reward bold well-designed tests regardless of outcome, and treat a disproven assumption as a save (a bad change caught before full rollout). Detaching ego and status from the win/loss outcome is what frees teams to test the big, uncertain ideas where the real wins hide.

Claim: With only about a third of experiments winning, treating losses as failures suppresses the bold tests that produce the biggest wins; high-performing programs reframe non-winners as learning and as prevented losses. Source: Kohavi et al. — experiment win-rate reality. Context: Reward the quality of the hypothesis and the test, not the outcome — or the team will only test safe, trivial changes.

RGM EXPERT TRICK
Crown a ‘best loss of the quarter’ to make celebrating losses real

Every leader says they celebrate learning from losses. Almost none do anything that makes it true, so teams correctly read the real incentive — don’t lose — and test timidly.

I make it concrete with a ritual: a ‘best loss of the quarter’ award for the well-designed test whose surprising negative result taught the most or prevented the costliest mistake. The team presents it; leadership applauds it; it goes in the deck.

Once a loss can win an award, the fear that drives safe, trivial testing evaporates — and the bold tests where real wins hide finally get proposed.

WHY IT’S RARE · Talking about celebrating losses changes nothing; ritualizing it does. A visible ‘best loss’ award is the cheapest, most powerful signal that bold, well-run tests are safe to run.

Skills and team

A program needs a blend of skills, not just a tool admin: research and analytics (to find what to test and read results honestly), statistical literacy (to avoid fooling yourselves), UX/design and engineering (to build trustworthy variants), and program management (to run the backlog and cadence). On small teams one person wears several hats; as you scale, a center-of-excellence model — a core team that sets standards, trains, and reviews while embedded teams execute — is the common way to keep quality high as volume grows.

Scaling from one team to the enterprise

Scaling experimentation is a quality-control problem: as more teams self-serve, the risk isn’t too few tests but too many bad ones — under-powered, peeked-at, SRM-broken — eroding trust in the whole program. Mature orgs scale with guardrails: shared tooling with automated validity checks (SRM, sample-size enforcement), an education program so every tester knows the unforgiving rules, a review or certification step for high-stakes tests, and a central team that sets standards. Velocity without trustworthiness is worse than a small, rigorous program.

What’s the biggest barrier to experimentation culture?
The HiPPO — the Highest-Paid Person’s Opinion overriding test results. If evidence loses to seniority, no tool or process matters. Executive sponsorship and a transparent, evidence-ordered backlog are the standard defenses.
How do I get leadership to support experimentation?
Convert one influential leader with a high-stakes surprise: test something they were sure about and let a counterintuitive, money-saving result make the case. Then make that leader the visible champion of ‘we test, and we follow the test.’
Why celebrate losing tests?
Because ~2/3 of ideas don’t win, and punishing losses makes teams test only safe, trivial changes. Celebrating the learning (and the bad changes prevented) frees teams to test the bold, uncertain ideas where the biggest wins are.

Where culture fails

Experimentation cultures fail when: the HiPPO overrides results, leadership claims to be data-driven but never lets evidence change its mind, losing tests are punished (so teams play it safe), testing is harder than not-testing (so people skip it), the org chases test volume without quality guardrails, and there’s no shared home for learnings so the same lessons get re-learned. All are organizational, and all cap the technical program below its potential.

HiPPO overrides results

Seniority beating evidence teaches everyone that testing is theater.

THE MOVE · Secure a sponsor who follows the test publicly; order the backlog by evidence, not rank.
Claiming data-driven, acting otherwise

Leaders who never let data change their minds have a culture problem dressed as a tooling one.

THE MOVE · Make ‘has evidence ever changed our decision’ a real, answered question.
Punishing losing tests

Penalizing losses pushes teams to safe, trivial, unlosable (and unteachable) tests.

THE MOVE · Reward hypothesis and test quality, not outcomes; celebrate learnings and prevented losses.
Testing harder than not testing

If the rigorous path has more friction, people route around it.

THE MOVE · Engineer defaults so shipping behind an experiment is the easy path.
Volume without guardrails

Scaling self-serve testing without quality control floods the org with untrustworthy results.

THE MOVE · Scale with automated validity checks, education, review, and a standards-setting core team.

Your culture checklist

Culture is the rate limit — and it’s buildable. Tick what is genuinely true of your organization.

The operating checklist — tick what is true today
Scored. Progress saves on this device.0/10
CASE-method test

Prove it. Earn your passcode.

Ten questions, CASE method (Context · Analysis · Strategy · Execution). Pass at 90% to unlock this module’s completion passcode — retake as many times as you like.