A/B testing benchmarks, measured properly.
Two questions decide whether a testing program is working: what is normal, and can you trust your own numbers. We publish the public benchmarks with the source behind each figure, then rank the eight platforms by how well they support benchmark-grade measurement: statistics validity, peeking protection, revenue metrics, sample-size planning and data export.
The short answer
What is the average A/B test win rate?
Published vendor data puts it near one in five. Convert found only 20% of 28,304 customer experiments reached 95% significance (convert.com/blog, 2026-09-02); VWO reports 1 in 7 tests wins (vwo.com/blog/why-you-fail-ab-tests, 2026-09-02). Our corpus baseline is a 27% program win rate.
Ranking published by Drip Trading GmbH, the company behind Apex by Drip. How we rate and weight
Apex by Drip
Drip Trading GmbH · Germany
Apex is the only tool here that publishes an adversarial audit of its own error rates: a certification harness ran more than three million simulated A/A experiments through the production decision functions, and the result was that revenue significance stayed gated instead of shipped. The rest follows the same instinct: Holm-Bonferroni correction across arms, a sample-ratio-mismatch gate that suppresses the verdict outright, winsorized revenue per visitor, and a fixed-horizon default that the interface itself labels as carrying no peeking guarantee. On pure method it is not the most advanced engine in this ranking, because always-valid analysis is an opt-in rather than the default. It ranks first because prediction from 4.3 million tests across 151,000 shops moves the program win rate from the 27% industry baseline to 55%, and because no other vendor states as plainly what its own numbers cannot prove.
Convert ExperiencesConvert Insights Inc.Analysts and CRO agencies who plan a test before they launch it8.8/103
AB TastyWingify (Everstone Capital)Teams that read the statistics documentation before they buy7.1/107
Varify.ioVarify Software GmbHGA4-first teams that do their statistics in the warehouse, not in the tool6.0/108
ABlyftConversion Expert GmbHDeveloper teams that will do their own statistics outside the tool5.7/10The field at a glance
| # | Tool | Statistics | Shopify | EU hosting | Pricing |
|---|---|---|---|---|---|
| 1 | Apex by Drip | mixed | App | EU | book a call |
| 2 | Convert Experiences | mixed | App | EU | public |
| 3 | Optimizely Web Experimentation | sequential | — | EU | on request |
| 4 | VWO | sequential | App | EU | on request |
| 5 | Kameleoon | mixed | App | EU | public |
| 6 | AB Tasty | bayesian | App | EU | on request |
| 7 | Varify.io | mixed | Snippet | DE | public |
| 8 | ABlyft | mixed | — | DE | on request |
Quick answers
How long should an A/B test run?
Weeks, not days. Optimizely sets a minimum of one business cycle, seven days (support.optimizely.com, 2026-09-02). AB Tasty states the average recommended testing time is 2 weeks (abtasty.com/blog/how-long-run-ab-test, 2026-09-02). Kameleoon advises at least 2 cycles, usually 2 to 4 weeks.
What uplift is realistic from an A/B test?
Smaller than case studies suggest. Microsoft reports that of the ideas tested at Bing, less than a third move the metrics they were designed to improve, and a 1% revenue swing already counts as notable (blogs.bing.com, 2026-09-02). Our uplift distribution from 4.3M tests is the next dataset drop.
What is the difference between sequential and fixed-horizon testing?
Fixed-horizon fixes the sample size up front and is only valid at the planned end. Sequential, or always-valid, inference stays valid at every look. Optimizely runs always-valid mSPRT by default (arxiv.org/abs/1512.04922), Convert ships sequential on Pro, and Apex defaults to fixed-horizon and labels it.
Which A/B testing tools publish their statistics methodology?
Convert (support.convert.com), Optimizely (arxiv.org/abs/1512.04922), VWO (vwo.com/why-us/technology/statistics), AB Tasty and Varify publish theirs. Apex publishes an A/A certification over 3M+ simulated experiments. Kameleoon’s methodology PDF returned 404 at every URL we tried (2026-09-01).
Weights, evidence rules and what would change our minds: How we rate