Blog

How to A/B test lead forms (without fooling yourself)

A practical guide to testing lead forms properly: what to test, how many conversions you need, and the traps that make results lie to you.

Marian Leonte, founder12 min read

Running an A/B test on a lead form is easy. Running one whose result you can actually trust is the hard part. Many "we tested it and conversion went up 20%" claims you'll hear from media buyers would not survive contact with a calculator: the sample was too small, the traffic mix shifted mid-test, or the win only holds for the first three days before it fades back to nothing.

This is a guide to doing it properly: what's actually worth testing on a form, how to set up a test that isolates one variable, how many conversions you need before you believe the result, the traps that quietly wreck an otherwise clean test, and how to check that a "win" on completion rate isn't secretly a loss on lead quality. There's a short section at the end on doing this in Formglide specifically, but everything before it applies whichever tool you use.

What's actually worth testing on a form

Not every change is worth burning traffic on. Some elements move completion rate a lot; others are cosmetic. Roughly in order of expected impact for a paid-traffic lead form:

What to test Why it tends to move the needle Typical risk if you get it wrong
Question count and order Fewer questions before the "yes" moment usually raises completion; but cutting a qualifying question raises completion while lowering lead quality Optimising for volume, not revenue
One-question-at-a-time vs. multi-question page Conversational, single-question flows often feel lower-friction; a single scrollable page can feel faster for people who know what they want Depends heavily on traffic source and intent: test, don't assume
The first question It sets the commitment level; an easy, low-stakes first question (e.g. "What's your goal?") tends to raise starts more than a hard one (e.g. "What's your budget?") A first question that filters too early can suppress volume you actually wanted
Qualifying questions (budget, timeline, company size) Cuts completion rate almost every time, but improves the quality of who's left Removing them to "win" completion rate is the single most common way people fool themselves: more on this below
CTA copy on the final step Small effect on its own, but nearly free to test Rarely worth more than one round of testing before diminishing returns
Social proof on the form (logos, review counts, response numbers) Can lift completion, especially on cold traffic that doesn't know your brand Needs a genuine claim behind it: fabricated numbers are a legal and trust problem, not just a marketing one

That ordering is a rule of thumb from practice, not a guarantee: your traffic and offer may behave differently, which is the whole reason to test. Test the things nearer the top of that list first. A CTA-copy test on a form that's still asking eight questions before the useful one is testing the wrong variable.

How to set up a test you can actually trust

Four rules, all of which are more often broken than followed:

Change one thing at a time. If variant B has a shorter question list and different CTA copy, a win tells you the combination worked, not which part did. If you want to know which lever matters, isolate it.

Split traffic 50/50 with a real assignment mechanism, not a schedule (e.g. "variant A this week, variant B next week"). Scheduling folds day-of-week and seasonality effects into your result and makes it useless. A proper split assigns each visitor to a variant at random, consistently, for the duration of their session: normally done with a cookie so a returning visitor sees the same variant they started with.

Don't touch spend or targeting mid-test. If you raise budget, add an audience, or swap creative halfway through, you've changed the input mix feeding both variants and you can no longer tell whether a shift in results is the form or the traffic.

Run full weeks, not partial ones. Weekday and weekend audiences behave differently for most lead-gen categories: B2B traffic is heavier midweek, consumer traffic often skews weekend. A test that runs Tuesday to Thursday isn't a sample of your traffic, it's a sample of your Tuesday-to-Thursday traffic. Run in multiples of seven days.

How many conversions you need before you believe the result

This is the step almost everyone skips, and it's the reason so many "winning" tests stop winning once you scale them.

The plain-English version: a small sample can produce a big-looking difference by chance alone. If variant A gets 3 conversions from 10 visitors and variant B gets 5 from 10, B looks 67% better, but with numbers that small, that gap is well within the range you'd expect from coin-flip noise, not a real effect. The fix isn't a special trick; it's comparing your two conversion rates with a standard two-proportion significance test, which tells you how likely it is that a gap this size would show up even if both variants were identical.

You don't need to run the maths by hand. A sample-size calculator built on the same two-proportion test (Evan Miller's A/B test sample size calculator is a widely used, free one) will tell you, before you launch, how many visitors each variant needs to reliably detect the lift you're hoping for.

A worked example. Say your current form completes at 30%, and you want to know if a shorter version can lift that to 34% (a 4-percentage-point, or roughly 13% relative, improvement). Using the standard formula behind that kind of calculator (comparing two proportions at 95% confidence and 80% power, the conventional defaults for this kind of test), that works out to roughly 2,130 visitors per variant who reach the form, around 4,260 in total, before you can trust a result of that size. At 30-34% completion, that's only about 640-720 completed forms per variant, but it's the visitor count that sets how long the test has to run.

Two things fall out of that example that are worth internalising:

  • Smaller expected lifts need bigger samples. Detecting a 1-point lift needs a much larger sample than detecting a 4-point one, for the same confidence level: the calculator will show you how fast the number grows as your expected effect shrinks.
  • Low-traffic accounts should test bigger, more obvious changes (cutting three questions, not rewording a button) rather than small tweaks they'll never collect enough data to confirm either way.

If your account doesn't move enough monthly volume to hit the sample size a calculator gives you within a reasonable time, that's useful information too: it tells you to test bigger swings less often, rather than small ones you can never actually resolve.

Turning a sample size into a calendar. Divide the total visitors you need by your daily visitors to the form, then round up to full weeks. In the example above, 4,260 visitors at 310 form visitors a day is just under 14 days: two full weeks. At 60 visitors a day it is 71 days, which rounds up to eleven weeks: at that point the right move is usually a bigger change, a bigger audience for the test, or no test at all and a decision made on judgement. A test you cannot finish is not a cheap test, it is a slow way to learn nothing.

Write the decision rule down before launch. For example: "Ship B if it beats A on completion rate at the planned sample size and the share of leads that reach a booked call is not lower. If neither condition is met, keep A." Deciding in advance removes the temptation to rationalise a result you were hoping for.

Traps that quietly wreck a "clean" test

Peeking. Checking the result every day and stopping the moment it looks significant inflates your false-positive rate dramatically: in Evan Miller's analysis, a test meant to run at a 5% false-positive rate can produce false positives about 26% of the time in the worst case, where you check after every new visitor and stop the moment it looks significant, and checking an experiment ten times turns what looks like 1% significance into roughly 5% (Evan Miller, "How Not To Run An A/B Test"). Decide your sample size before you start, and don't call the result until you hit it.

Novelty effects. A redesigned form can spike simply because it's different, not because it's better: the lift shows up in the numbers but wears off, and it would be wrong to credit it to the change itself (Analytics Toolkit glossary: novelty effect). This matters most for visually dramatic changes. The guard against it is the same one as against peeking: run longer, especially for bigger visual changes, so the initial reaction has time to wear off before you call a winner.

Mixing traffic sources mid-test. If variant A ran only against Meta traffic and variant B picked up a chunk of Google traffic when you added a new campaign, you're not comparing forms anymore: you're comparing traffic sources that happen to be looking at different forms. Keep the traffic mix constant, or split by source and analyse separately.

Optimising completion rate while lead quality quietly drops. This is the trap that costs the most money and shows up last. Cut a qualifying question and completion rate goes up: of course it does, you removed friction. But if that question was filtering out people who were never going to buy, you've just paid to collect more leads that convert worse downstream. A form-level "win" that isn't checked against what happens after the form isn't a win, it's a shifted problem.

Measuring lead quality, not just completion rate

The fix for the last trap is to follow leads past the form, not just to it.

Tag every submission with its source. UTM parameters and hidden fields captured on submission (campaign, ad set, creative, landing page) let you trace a completed form back to the traffic that produced it: without that link, you can't tell whether a completion-rate win came from better traffic or a better form.

Track leads through your CRM stage, not just form completions. A lead-quality comparison worth trusting looks at what proportion of each variant's leads reached a qualified, meeting-booked, or closed stage in your CRM: not just how many filled in the form. Pull this by variant (which you can do if your form or CRM records which variant produced each lead) over the same window you ran the form test, and compare qualified-lead rate alongside completion rate, not instead of it.

Give it time to show up. Completion rate is visible the day the test ends; sales-qualified rate might take weeks to resolve, depending on your sales cycle. Don't declare a form change a win until you've checked both.

Doing this in Formglide

Formglide's A/B testing is on the Pro plan and above. This is how it works today, including the parts that are limited:

  • How you set up B. Variant A is the version of the form that is currently published. Variant B is whatever is in your unpublished draft at the moment you press "Start A/B test" in the form's settings. So make your change in the draft first, then start the test. If you start a test with no draft changes, B is an identical copy of A and the test tells you nothing.
  • The split. New visitors are split 50/50 at random. The assignment is stored in a cookie (fg_ab, one year), so a returning visitor keeps seeing the variant they started with. The split is fixed at 50/50; there is no weighted split.
  • What you see. The A/B report on the form's analytics page shows views, completions and completion rate for each variant, plus B's percentage difference from A. Under 100 completions, a variant is labelled "Early signal". The numbers cover the date range picked on that page (7, 30 or 90 days), so choose one that spans the whole test.
  • What you do not get: significance. The report is explicitly a naive comparison. It does not calculate statistical significance, so it will not tell you whether a gap is real or noise. Take the views and completions per variant from the report, run them through a two-proportion significance calculator, and hold yourself to the sample size you calculated before launch. The 100-completion label is a guard rail, not a verdict: 100 completions per variant is nowhere near enough to detect a small lift.
  • Ending the test. You can pick A or B as the winner, which makes that variant live for everyone, or stop the test. Only one test can run on a form at a time.
  • Lead quality by variant. The CSV export (available on every plan) includes a variant column next to the utm_source, utm_medium, utm_campaign, utm_term and utm_content columns, so you can join submissions to your CRM and compare qualified-lead rate by variant. UTM parameters are picked up from the form's URL automatically; add any hidden fields you want to carry (an ad ID, a landing-page name) before the test starts, not after. One catch: while a test runs, the export's answer columns follow the published version (A), so a question that exists only in B won't get its own column.
  • Where it sits in the funnel. Per-step drop-off analytics (also Pro) show which step the published form loses people on. The funnel isn't split by variant yet, so judge a length test on the A/B report's completion numbers and use the funnel to find the leak worth testing next.

None of this is unique to Formglide: you can run the same disciplined test in any tool that splits traffic randomly and reports per-variant numbers. What Formglide tries to do is keep the split, the UTM capture and the drop-off view in one place instead of stitched together from three tools. If you want to try it, there is a free plan to build the form on and Pro is priced on the pricing page.

A pre-launch checklist

  • One variable changed between variants, not several at once
  • Traffic split 50/50 by a consistent, cookie-based assignment
  • Ad spend and targeting locked for the duration of the test
  • Test scheduled to run in full weeks (7, 14, 21 days), not partial ones
  • Sample size calculated in advance, based on your current completion rate and the lift you want to detect
  • A decision made not to look at significance until the sample size is hit, and a written rule for what counts as a win
  • UTM or hidden-field tagging in place, and the variant column kept in your export, so you can trace leads back to their variant
  • A plan to check CRM-stage lead quality by variant, not just completion rate, before declaring a winner

Whichever tool you're using, the discipline is the same: one change, a real split, a sample size decided before you start, and a check on what happens to the leads after the form, not just at it.

Build your first form free

No credit card. 500 responses/month, forever, never shuts off.