Attach a testing tool and it will split traffic between two screens and show you numbers. The tool is not the hard part. The hard part is when you decide when to stop.
Leave a test running and watch the dashboard daily, and early on the two figures swing wildly. Then one morning your variant B is ahead. Stopping there records the moment chance happened to favour you as a result. If the gap evaporates next week, you have already shipped.
The two decisions that come first
Sample size answers “how many people before I can judge”. The lower your conversion rate, and the smaller the difference you want to detect, the more visitors you need. Calculators for this are freely available: put in your current rate and the difference that would be worth acting on, and use the number that comes out.
Duration is not set by sample alone. Cover whole cycles, at least full weeks. Weekday and weekend visitors behave differently, and campaign bursts or newsletter sends cluster on particular days. A test that ran Wednesday to Friday may have measured “Thursday traffic is different” rather than “B is better”.
Write it down before you start
There is one primary metric. Watching click-through, form submissions and final conversion together and then announcing that one of them rose is choosing which roll of the dice to keep. Record the rest as context, but do not decide on them.
And the hypothesis must be one sentence: “rewriting the hero line around outcomes will increase form submissions”. If variant B changed the copy, the image and the button all at once and won, what you learned is that the combination is better — not why.
What if you do not have the traffic?
Honestly: many sites do not have enough traffic for an A/B test to ever conclude. The answer is neither to extend the run indefinitely nor to call it early on a short sample. It is to use a different method.
Watch where people actually stall, read the form error logs, collect the questions arriving through enquiries, and walk the path from entrance to completion yourself. Qualitative work does not hand you a percentage, but it tells you what to fix — and fixing something visibly broken never needed an experiment.
Finally, check what the tool itself costs. Client-side testing tools often intervene before the page paints, so switching one on can delay the first screen. Compare medians of three runs on the same URL either side of installing it, and you will not spend a fortnight hunting that cause later. Turning experiments into a repeatable procedure is covered in the development workflow archive and on our process page.
Next part
The last part is the conversion audit: before experimenting on individual elements, finding where people leave along the whole path from entrance to completion.