57% of A/B test winners wouldn't hold up at proper sample size.
Most testing programs stop the moment a result looks good, not when it's actually proven. We design, run, and analyze tests built on real statistical rigor — sample size calculated first, significance checked before anything ships.
What do A/B testing services include?
A/B testing services include hypothesis formation, sample size calculation, test design, statistical analysis, and program management for controlled experiments on a website, app, or email — covering the statistical rigor behind a test rather than just the visual variant a designer builds.
The uncomfortable number every testing program should know: 57% of A/B tests called winners would not have reached statistical significance if run to their proper sample size. In a separate analysis of 28,304 experiments, only 20% reached the standard 95% confidence threshold. Most “wins” teams ship never actually happened — the test just got called too early.
- Sample size first
- Calculated before launch, so the stop date is set in advance.
- No peeking
- Results are not read early, because that invalidates the maths.
- Full weekly cycle
- 2–6 weeks, since weekday and weekend traffic differ.
- SRM checked
- Sample ratio mismatch caught before any result is called.
A/B testing, CRO, or landing page optimization?
All three sit in the same workflow. Each does a different job in it.
| Service | Handles | Best fit |
|---|---|---|
| A/B Testing | Sample size, test design, and statistical proof that a change actually works. | Validating a specific hypothesis with real confidence, not a guess. |
| Conversion Rate Optimization | What to test and why, based on user research and funnel data. | Building the overall testing roadmap and priority order. |
| Landing Page Optimization | Full-page implementation — copy, design, and layout together. | Rebuilding a page once winning variants are proven. |
Smaller lifts need dramatically more traffic to prove.
Which is why low-traffic sites should test bold changes, not subtle ones.
The sample size problem, on a 5% baseline
Detecting a 20% relative lift takes roughly 15,000 visitors per variation. Detecting a 5% relative lift on the same baseline takes roughly 240,000. That 16-fold jump in required traffic to catch a smaller effect is why testing button colors on a page with limited traffic almost never reaches a valid sample, no matter how long the test runs.
What actually invalidates a test
Peeking — stopping the moment a result looks good, before the planned sample is reached. Winner's curse — apparent lift shrinking once the full sample is in. SRM — sample ratio mismatch, where an unequal traffic split silently skews results. And fixed days — running by the calendar instead of the calculated required sample.
A/B testing is the proof layer underneath CRO strategy. Without it, a CRO roadmap is just a list of opinions.
A/B testing work businesses bring us.
The statistical layer, not just the variant a designer builds.
Hypothesis & test design
We turn a data-backed observation into a testable hypothesis with a clearly defined success metric.
Sample size & duration planning
We calculate required visitors and test duration before launch, so the stop date is set in advance.
Test implementation
Variant build and QA across devices, so the only difference between A and B is the one being tested.
Statistical analysis
Significance, confidence intervals, and sample ratio checks run before any result gets called a winner.
Multivariate testing
Testing multiple elements at once when traffic supports it, without sacrificing statistical validity.
Experimentation program management
An ongoing test roadmap and backlog, so wins compound quarter over quarter instead of running as one-offs.
A clear path from hypothesis to a verified result.
Four stages, with the stop date locked before the test ever goes live.
Hypothesis & baseline data
We pull the current conversion rate and define the minimum lift worth acting on.
1 week · PlanningSample size calculation
We calculate the required sample and lock the test duration before anything goes live.
2–3 days · CalculationRun to full sample
We run the test to its planned sample size — no early stops, no peeking at daily results.
2–6 weeks · ExecutionAnalyze & report
We report significance, confidence interval, and practical impact, then queue the next test.
1 week · AnalysisThe tools we use for test design and analysis.
Testing platforms plus the statistical and behavioural tooling around them.
Explore more Conversion Optimization services.
A/B testing is one of three services we offer under Conversion Rate Optimization.
Conversion Rate Optimization
The strategy layer above testing.
ExploreLanding Page Optimization
Full-page CRO once tests are proven.
ExploreUX Audit
Where the hypotheses worth testing come from.
ExploreCopywriting
The variant copy going into every test.
ExploreKlaviyo Email Marketing
Where flow & subject line tests run.
ExploreGoogle Ads Management
Ad creative tested with the same rigor.
ExploreBusiness Intelligence
The data layer behind every baseline.
ExploreBranding
The identity constraints every variant works inside.
ExploreA/B testing questions
The things clients ask us most before starting a testing program.
A/B testing services include hypothesis formation, sample size calculation, test design, statistical analysis, and program management for controlled experiments on a website, app, or email — covering the statistical rigor behind a test rather than just the visual variant a designer builds.
No, not without proper sample size discipline. Research shows 57% of A/B tests called winners would not have reached statistical significance if run to their proper sample size, and in a separate analysis of 28,304 experiments, only 20% reached the standard 95% significance threshold. Most test failures come from stopping early, not from a bad idea.
It depends entirely on the size of the improvement being detected. Detecting a 20% relative improvement on a 5% baseline conversion rate requires roughly 15,000 visitors per variation, while detecting a smaller 5% relative improvement on the same baseline requires roughly 240,000 visitors per variation — a 16-fold jump in required traffic for a much smaller effect.
Statistical significance is the confidence that a result is caused by a real effect rather than random chance. Most conversion rate optimization programs use a 95% confidence threshold, meaning there's a 5% probability the observed difference happened by chance rather than because the variant genuinely performed better.
The peeking problem is checking test results before reaching the predetermined sample size and stopping as soon as significance appears, which invalidates the statistical calculation and sharply increases the false positive rate. A test that looks significant on day 4 often reverts once it runs to its full planned sample.
Conversion rate optimization is the broader strategy — deciding what to test and why, based on user research and funnel data. A/B testing is the statistical execution layer underneath it — sample size, test design, and significance analysis — that determines whether a CRO hypothesis actually gets proven or just guessed at.
Generally 2 to 6 weeks, long enough to reach the pre-calculated sample size and to capture a full weekly business cycle, since weekday and weekend traffic often behave differently. Running a test for a fixed number of days instead of a calculated sample size is one of the most common ways teams invalidate their own results.
Yes. About 63% of businesses running a systematic, ongoing A/B testing program report measurable revenue increases, since gains compound: a verified lift this quarter stacks on top of last quarter's verified lift, rather than each change being a one-off guess.