Refiner

a/b testing agency

A/B Testing & Experimentation

An a/b testing agency that runs statistically sound experiments on landing pages, email and paid creative — built for B2B, tech and fintech sales cycles.

Most B2B teams that say they're 'testing' are actually just trying things sequentially and comparing the results informally afterwards — which isn't testing, it's guessing with extra steps. Working with an a/b testing agency should mean something more specific: properly randomised experiments, run for long enough to reach statistical confidence, on changes big enough to plausibly matter.

We design and run experimentation programmes for tech, fintech and professional services companies where traffic and lead volume are often lower than the ecommerce world that most testing methodology was originally built for. That constraint changes almost everything about how a B2B testing programme should be designed, and it's the first thing we address before running a single test.

10–25%

Typical conversion lift on winning tests

4–8

Tests run per quarter, active programme

3–6 weeks

Time to statistically usable result

Why B2B testing needs a different playbook

Standard A/B testing methodology assumes traffic volumes that most B2B companies simply don't have. An ecommerce site can reach statistical significance on a button colour change within days; a B2B SaaS company with 2,000 monthly landing page visitors and a 3% conversion rate might need months to reach the same confidence on a change of similar size.

The practical response isn't to abandon testing — it's to test bigger, more consequential changes less frequently, and to use sequential and Bayesian methods that give usable directional confidence sooner than strict frequentist significance thresholds allow, while being honest about the uncertainty that remains. We're explicit with every client about what confidence level a given test can realistically reach given their actual traffic, before the test launches, not after.

We also lean more heavily on qualitative and behavioural signals — session recordings, funnel drop-off points, sales call feedback about objections — to generate test hypotheses worth the traffic cost of running, rather than testing changes that look interesting on a competitor's blog post but have no evidence behind them for your specific audience.

What we actually test

The highest-leverage tests in B2B marketing are rarely cosmetic. Headline and value proposition framing on primary landing pages, the structure and length of demo or trial request forms, pricing page layout and packaging presentation, and email subject lines and send cadence in nurture sequences consistently produce larger, more reliable effects than micro-level design tweaks.

For fintech and financial services clients specifically, we frequently test how risk, compliance and security messaging is framed and positioned — because for regulated buyers, reassurance on this front often matters as much as the core value proposition, and getting the framing wrong can quietly suppress conversion in ways that are easy to miss without deliberate testing.

For professional services firms, where a single landing page might only get a few hundred visits a month, we shift a larger share of the experimentation budget towards email and outbound messaging testing, where volume is higher and iteration is faster, while still running slower, longer-duration tests on the small number of high-traffic pages that do exist.

Running a test properly

Every test starts with a written hypothesis, not just a variant to try: what we believe is happening, why, what change we're making, and what result would confirm or disprove the hypothesis. This discipline matters because it stops teams from quietly reframing a failed test as a success after the fact by finding a metric that happened to move favourably.

We calculate the minimum sample size and expected test duration before launch, based on your actual current conversion rate and traffic, and we commit to that duration in advance rather than checking results daily and stopping the moment a result looks favourable — a common and statistically invalid practice known as peeking, which dramatically inflates the rate of false positives.

Where sample sizes are genuinely too small to support a rigorous split test, we say so directly and recommend a phased rollout with careful before-and-after measurement instead, which is a legitimate and often underused alternative for lower-traffic B2B pages rather than forcing a test methodology that the traffic can't support.

Interpreting results honestly

A test result is a data point, not a verdict on your entire strategy. We report results with confidence intervals, not just a single 'winning' percentage, and we're explicit about the difference between statistical significance and practical significance — a result can be statistically real and still too small to be worth the operational cost of implementing permanently.

We also track whether winning variants continue to perform after being rolled out fully, because novelty effects and seasonal factors can make a test result look stronger during the test window than it holds up to afterwards. This follow-up check is skipped by most testing programmes we've inherited from other agencies, and it's often where we find that a previous 'win' had already faded.

Building an experimentation backlog

Rather than testing reactively, we build a prioritised backlog of hypotheses scored on expected impact, confidence in the underlying hypothesis, and ease of implementation — a standard ICE-style framework, applied consistently so that the loudest voice in the room doesn't automatically win the next test slot.

This backlog is reviewed monthly and refreshed with new hypotheses generated from the growth audit findings, sales team feedback, customer research and previous test results, so the programme compounds — each test informing the next rather than existing as an isolated one-off experiment.

Where testing fits inside compliance constraints

For regulated fintech clients, we agree an approved range of messaging variants with compliance and legal before testing begins, covering the categories most likely to need review — claims about returns, risk, regulatory status and security — so that individual test variants within that pre-approved range don't each require separate sign-off, which is usually the biggest blocker to running experimentation at any meaningful pace in regulated environments.

Frequently asked

We don't get much website traffic — is A/B testing still worth it?

It depends on the page and the size of change. Low-traffic pages need bigger, more consequential test variants and longer run times to reach usable confidence, and for very low-traffic pages a phased rollout with careful measurement is often more appropriate than a formal split test. We'll tell you honestly which approach fits your actual numbers before recommending either.

How do you decide what to test first?

We build a prioritised backlog scored on expected impact, our confidence in the underlying hypothesis, and implementation effort, drawing on growth audit findings, session recordings, sales objection data and previous test results. This keeps prioritisation evidence-led rather than defaulting to whatever change is easiest to build or most recently suggested in a meeting.

Can you run tests within our existing compliance and legal review process?

Yes — for regulated fintech and financial services clients we build a pre-approved messaging framework with your compliance team upfront, defining acceptable variant ranges for claims, risk and security language. This lets us run experiments within that approved range without each individual variant needing separate review, which is usually the main constraint on testing pace in regulated firms.

What tools do you use to run tests?

We work within whatever testing and analytics stack you already have where it's fit for purpose, and recommend specific tools when there's a genuine gap. The methodology and statistical rigour matter more than the specific platform, and we've run sound experimentation programmes on relatively simple toolsets when the underlying test design and measurement discipline were solid.

Refinement consultation

Let's talk a/b testing & experimentation

Answer four quick questions and we'll come back within one working day with a specific, costed way forward.

See the work