Skip to content

Article

July 28, 2026

8 min read

The A/B Test That Lies

A/B tests can establish causal movement, but teams still need a metric hierarchy that distinguishes local engagement from durable product trust.

By Cristiano Pierry

The A/B Test That Lies

Picture a familiar experiment readout. Clicks are up. Engagement is up. The confidence interval is clean. Every important cell is green, and the room relaxes.

The variant ships.

Two weeks later, the product feels worse. People are clicking, but some of those clicks look more like interruption than interest. Sessions are longer, but not necessarily better. The experience has become more aggressive, repetitive, or indifferent to why the user came.

Nothing in the readout was fabricated. The randomization worked. The instrumentation passed. The variant caused more of the behavior the team measured.

The test told the truth. The organization asked it to certify the wrong win.

Properly designed randomized tests can tell us whether a change caused an observed difference rather than merely appearing alongside it. They can challenge executive intuition, expose weak ideas, and protect teams from shipping the most persuasive opinion in the room.

A test can establish that a variant increased clicks. By itself, it cannot establish that the clicks were useful, the experience became more satisfying, or the product strengthened a relationship the company will still value six months later.

The statistical question may be settled while the management question remains open.

I recently argued that product teams need fewer dashboards and more dials. Experimentation is one of the most powerful dials a team can have: it changes the product, creates evidence, and gives the organization a disciplined way to learn. Connected to the wrong outcome, it can also help the team steer precisely in the wrong direction.

When experimentation becomes a slot machine

A team runs many tests. Each one produces a large readout. The organization scans for green. A lift appears somewhere important enough to tell a good story, the team celebrates the win, and the next test enters the queue.

Nobody has to manipulate the data for this system to drift. If teams earn recognition for experiment wins, they will favor changes that move sensitive, immediate metrics. Those metrics respond quickly, reach significance sooner, and fit neatly into weekly reporting.

The slower consequences arrive later and belong to everyone: fatigue, regret, reduced confidence, weaker retention, or a growing sense that the product is no longer acting in the user’s interest.

Steven Kerr described the management problem fifty years ago as the folly of rewarding A while hoping for B. Product organizations say they want durable customer value, then praise a local lift that can be observed before the next planning review. The experiment did not create that conflict. It made the reward system look scientific.

Microsoft’s research on long-term experimentation provides a useful example. Degrading search quality can increase search activity in the short term because users run more queries when the first results do not work. A dashboard watching query volume may register more engagement precisely because the experience became less successful.

The metric moved exactly as instrumented. The team still has to decide whether the movement represents value or recovery from failure.

The Metric Ladder

I find it useful to place experiment metrics on a ladder with four rungs: movement, engagement, satisfaction, and trust. Lower rungs are often faster and more sensitive. They help diagnose what changed. The mistake is letting one rung impersonate the one above it.

Movement is the immediate response. The user clicked, opened, started, dismissed, or selected something. A larger card may earn more clicks because it communicates value clearly. It may also earn them because it occupies more space or makes the alternative difficult to see. The movement is real; its meaning still depends on the design.

Search makes the ambiguity especially clear. Research on search-result evaluation has documented “good abandonment”, where a person does not click because the results page already answered the question. A click-only metric can interpret a useful answer as failure and a forced extra step as success.

Engagement asks whether the movement turned into deeper use. The person watched, read, listened, completed, explored, returned, or spent more time with the experience.

More time can mean fascination or confusion. Another session can reflect loyalty or an unfinished task. Autoplay can increase consumption while reducing the user’s sense of control. The team still has to ask what the person received in exchange for the additional engagement.

Satisfaction asks whether the interaction did the job the user came to do.

This usually requires signals such as task completion, reformulation, dead ends, explicit feedback, and qualitative review of the journey. Satisfaction is harder to observe than a click because it lives in the relationship between intent and outcome.

A streaming session that ends after twenty minutes may be a failure because the viewer gave up, or a success because they found the short piece they wanted. A search session with no click may be unsuccessful, or it may have answered the question immediately. The event count does not know.

Trust is the judgment the user accumulates about the product over time. Will it respect my attention, remain dependable, and act in my interest after the novelty wears off?

Trust is important to name and easy to oversimplify. Retention, direct return, renewal, opt-outs, complaints, corrections, and longitudinal research can all provide evidence. None captures the whole thing. A team should be suspicious of any trust metric that becomes as easy to optimize as click-through rate.

The experiment result may live on one rung while the launch decision needs evidence from another.

A win needs permission

Not every experiment has to reach the trust rung before a team can ship.

If a team is testing the wording of a secondary control, movement and task completion may be enough. If it is changing the default that shapes what millions of people see, introducing a more interruptive notification, altering a recommendation objective, or making an AI system sound more certain, the acceptable evidence should be different. The potential harm and reversibility of the decision should determine how high the test has to climb.

Before the test launches, the team should be able to state three things:

  1. What immediate behavior do we expect to move?
  2. Why do we believe that movement will improve satisfaction or trust?
  3. What evidence would tell us that the local win came at their expense?

The second question is often missing. The plan names the primary metric and guardrails without making the causal belief explicit. Once the result arrives, the missing logic gets filled in after the fact. Clicks become evidence of relevance. Time becomes evidence of enjoyment. Return frequency becomes evidence of loyalty.

A stronger experiment contract forces the team to write the product hypothesis before it knows which story the dashboard will make convenient.

A hand-drawn product team defines behavior, the causal reason, and a guardrail on an experiment contract before viewing the green result.
Write the behavior, the causal belief, and the guardrail before the result makes one story convenient.

Guardrails need the same discipline. A crash-rate guardrail will not catch a relevance problem. Session length will not catch a loss of control. Seven-day retention may miss a change that slowly teaches users to expect less from the product.

Microsoft’s experimentation guidance separates local diagnostic metrics, overall evaluation criteria, data-quality checks, and guardrails because they answer different questions. Click-through rate can explain how a feature behaved without earning the authority to decide whether it should ship.

The metric hierarchy therefore carries an organizational decision: which signal informs the launch, which one constrains it, and who remains accountable when they disagree?

Proxies have to earn authority

Teams cannot run every experiment for six months. Product conditions change, tests interfere with one another, and long-running holdouts introduce their own biases. Waiting for perfect evidence would turn disciplined experimentation into paralysis.

The alternative is to earn confidence in the proxies the organization uses.

Google researchers working on recommender systems have studied near-term behaviors that predict how often users with the same initial visiting frequency return five months later. Research on search advertising has also shown how previous ad quality can affect a person’s later propensity to engage. In both cases, the short-term metric becomes more useful because researchers tested its connection to a longer-term outcome instead of inheriting that connection as folklore.

Teams can build that evidence through selective long-running holdouts, staged rollouts, post-launch reviews, cohort analysis, experience sampling, and periodic qualitative work. A low-risk, reversible change can ship with a lighter contract. A change to core ranking, monetization pressure, notification defaults, or an AI system’s confidence may deserve a longer soak period and stronger guardrails.

Some experiments can move quickly because prior work validated the proxy. Others should move slowly because the product is entering a new risk area or the proposed win depends on a behavior the organization has never connected to durable value.

Someone should still own the question after launch. If the local metric remains green while satisfaction weakens, the organization has learned that its metric ladder needs repair.

Taste still has a job

An article about the limits of A/B testing can easily slide into a defense of executive instinct. That would be the wrong lesson.

Taste is not a license to dismiss uncomfortable evidence. Strategy cannot rescue a preferred idea after it loses. If a leader can override any result by saying the product “feels worse,” the experiment becomes theater of a different kind.

Product judgment should enter before the result through the user promise, metric hierarchy, guardrails, and level of evidence the decision requires. It remains relevant afterward through the review of segments, unintended effects, qualitative evidence, and consequences the test window could not observe.

The test contributes causal evidence. Taste helps the team notice what the metric cannot represent cleanly. Strategy decides which behaviors are worth causing. Accountability stays with the people who ship the change. Those roles overlap, but they are not interchangeable.

This matters most in products built around discovery, recommendations, feeds, notifications, and AI assistance. These systems can become more effective at producing a measurable response while becoming less worthy of the user’s attention. No single experiment readout will protect the product from that trade.

When the dashboard turns green, I would change the first question in the room. Before asking, “Did the test win?” ask, “What did it win?”

Which metric would make your team proud this week but embarrassed six months from now?


This writing reflects my personal perspectives on product management, AI, and content discovery. It does not represent the official position of my employer or any affiliated organization.