An A/B test is how you make the market vote before you commit. Done with discipline, it separates the changes that genuinely move revenue from the ones that merely feel clever in a meeting. Done sloppily, it manufactures false confidence, and false confidence scales losses. We run testing as an evidence practice, not a tool subscription.
Where testing programs go to die
Peeking: stopping at the first happy number, which promotes noise to strategy. Underpowered tests on pages without the traffic to conclude. Trivia hypotheses: fifty shades of button while the offer goes unquestioned. And amnesia: results undocumented, so the same losers get retested every eighteen months under new managers. Our process exists to kill all four.
What we handle
- Hypothesis pipelines fed by research: recordings, customer language and funnel data, ranked by impact, evidence and effort
- Statistical design before launch: primary metric, minimum detectable effect, sample size and stop rules, pre-registered
- Execution across tools and stacks, with QA that catches flicker, tracking gaps and segment contamination
- Honest analysis: significance with context, segment effects checked without torture, revenue impact projected conservatively
- Implementation specs for winners so gains become permanent code, not permanent experiments
- The learning archive: every test’s hypothesis, result and decision, building a compounding model of your customer
What velocity should look like
Sized to traffic, honestly: high-volume sites can conclude several tests monthly; leaner sites should run fewer, bigger swings, and use evidence-led changes with before-and-after measurement where testing math cannot close. We would rather run four conclusive tests a quarter than twelve inconclusive ones.
What win rate should we expect?
Industry-wide, roughly one in four tests wins cleanly; research-fed pipelines do meaningfully better. But losers pay too: a documented loser is a customer insight and a bullet dodged at scale. The compounding asset is the learning, not the streak.
Can you test more than web pages?
Yes: emails, pricing presentation, ad-to-page journeys, checkout flows and onboarding sequences. Anywhere a measurable decision happens, the method applies.
The pre-registration ritual that keeps results honest
Before any test launches, its card is written and locked: hypothesis with the evidence behind it, primary metric, minimum detectable effect, required sample size, run length, and the decision rule for every outcome including “flat”. The card kills the post-hoc storytelling that plagues testing programs, where losers get reframed as segment wins and peeking gets rationalized. When the test ends, the card decides; opinions had their turn before launch.
Our last agency reported an 80% win rate. Possible?
Possible only through methods that manufacture wins: peeking, cherry-picked segments, or metrics chosen after results. Honest programs live near the industry’s one-in-four, improved by research quality. A high win rate is not a credential; it is usually a confession.
Hypotheses flow from Landing Page Optimization research, targets from Funnel Optimization, and creative variants from Meta Ads testing systems.
Geeks Digital