Technote

Development workflow Practical

Series Speed and conversion Part 7 of 8

An A/B test starts with two decisions

If you do not fix the sample size and the duration before you start, you will stop the test at the moment it flatters the variant you preferred. That is not an experiment.

Attach a testing tool and it will split traffic between two screens and show you numbers. The tool is not the hard part. The hard part is when you decide when to stop.

Leave a test running and watch the dashboard daily, and early on the two figures swing wildly. Then one morning your variant B is ahead. Stopping there records the moment chance happened to favour you as a result. If the gap evaporates next week, you have already shipped.

The two decisions that come first

Reorder this and it is no longer an experiment — the criteria come before the run

Sample size answers “how many people before I can judge”. The lower your conversion rate, and the smaller the difference you want to detect, the more visitors you need. Calculators for this are freely available: put in your current rate and the difference that would be worth acting on, and use the number that comes out.

Duration is not set by sample alone. Cover whole cycles, at least full weeks. Weekday and weekend visitors behave differently, and campaign bursts or newsletter sends cluster on particular days. A test that ran Wednesday to Friday may have measured “Thursday traffic is different” rather than “B is better”.

Write it down before you start

Decide these before running — the bottom two are ways of manufacturing a result

There is one primary metric. Watching click-through, form submissions and final conversion together and then announcing that one of them rose is choosing which roll of the dice to keep. Record the rest as context, but do not decide on them.

And the hypothesis must be one sentence: “rewriting the hero line around outcomes will increase form submissions”. If variant B changed the copy, the image and the button all at once and won, what you learned is that the combination is better — not why.

What if you do not have the traffic?

Honestly: many sites do not have enough traffic for an A/B test to ever conclude. The answer is neither to extend the run indefinitely nor to call it early on a short sample. It is to use a different method.

Watch where people actually stall, read the form error logs, collect the questions arriving through enquiries, and walk the path from entrance to completion yourself. Qualitative work does not hand you a percentage, but it tells you what to fix — and fixing something visibly broken never needed an experiment.

Finally, check what the tool itself costs. Client-side testing tools often intervene before the page paints, so switching one on can delay the first screen. Compare medians of three runs on the same URL either side of installing it, and you will not spend a fortnight hunting that cause later. Turning experiments into a repeatable procedure is covered in the development workflow archive and on our process page.

Next part

The last part is the conversion audit: before experimenting on individual elements, finding where people leave along the whole path from entrance to completion.

More on this topic

All technotes

Development workflow Advanced

The design QA checklist to run before release

Most design problems found after release are catchable before it, in order. Here is the pass, in three stages: rules, states and real devices.

Designers 6 min read

Development workflow Advanced

Never deploy without a way back

A rollback is not one button. It is two tracks — code and database — with an order to reverse them in, written down before you deploy.

Designers 9 min read

Development workflow Advanced

Adding your own hooks: extension points instead of edits

Code with no extension points gets forked or edited in place. Where you put do_action and apply_filters — and how many — decides how long that code survives.

Developers 7 min read

₩270,000 · Join the program