CRO, Experimentation & Personalization
Most testing programmes fail quietly: too few tests, too little traffic, tests stopped when they look good, and results nobody acts on. We build experimentation programmes that hold up — grounded in research, prioritised by expected value, powered correctly, analysed honestly, and connected to personalization so the wins compound instead of resetting.
Challenges we solve
The failure modes are consistent: tests launched on a hunch rather than research, sample sizes never calculated, experiments stopped the moment they cross significance, primary metrics chosen after the fact, and wins reported that never show up in the P&L. Then everyone concludes that testing does not work. A programme that works looks different. Research first, so hypotheses come from evidence. Prioritisation by expected value rather than by who asked loudest. Power calculated before launch and a fixed stopping rule. A single pre-declared primary metric with guardrails. And a searchable archive of what was learned, including the losses — which are usually where the expensive lessons are.
What we deliver
Conversion research and opportunity sizing
Quantitative analysis of funnels, segments and journeys combined with qualitative input — session replay, heuristic review, user testing, surveys and support-ticket analysis — to find where the money actually leaks and how much is at stake.
Hypothesis backlog and prioritisation
Structured hypotheses with a stated problem, proposed change, expected effect and reasoning, scored on expected value, confidence and effort so the roadmap reflects likely return rather than internal enthusiasm.
Experiment design and statistics
Power and sample-size calculation before launch, a single pre-declared primary metric with guardrail metrics, a fixed stopping rule, sequential testing where appropriate, and correction for multiple comparisons. This is the part that determines whether your results are real.
Test build and QA
Front-end implementation in Optimizely, VWO, Adobe Target or GrowthBook with cross-browser and device QA, flicker prevention, SPA-safe implementation, accessibility checks on variants, and performance impact measured rather than assumed.
Server-side and feature-flag experimentation
Moving testing out of the browser for pricing, algorithm, backend and app experiments — faster, flicker-free, and able to test things a client-side tool cannot reach.
Personalization and targeting
Rules-based and model-driven personalization across content, offers, products and messaging, built on real behavioural and first-party signals — always held to the same measurement standard as an A/B test rather than assumed to work.
Analysis, reporting and learning archive
Results analysed against the pre-declared plan, segment analysis flagged as exploratory rather than confirmatory, and every experiment archived with hypothesis, design, result and interpretation so the programme accumulates knowledge instead of repeating itself.
Programme operations and enablement
Test velocity planning, calendar and collision management, stakeholder communication, a decision log, and training so your team can run and interpret experiments without us.
Platforms & Tooling
- Optimizely
- VWO
- Adobe Target
- GrowthBook
- GA4
- Adobe Analytics
- BigQuery
- Server-Side GTM
- Hotjar
- Contentsquare
Process
Research
Quantitative and qualitative research, opportunity sizing.
Prioritise
Hypothesis backlog scored and sequenced with stakeholders.
Design & build
Powered experiment designs, built and QA-tested.
Analyse
Analysis against the pre-declared plan, archived with learning.
Review
Programme performance, roadmap reset, capability transfer.
FAQs
Then testing is the wrong first tool, and we will say so. Below meaningful sample sizes the better route is research-led change with careful before-and-after measurement, testing higher up the funnel where volume is greater, or using bandits for allocation rather than inference. Running underpowered tests and acting on the results is worse than not testing.
Until it reaches the sample size calculated before launch, and at minimum a full business cycle — usually two to four weeks — so you capture weekday and weekend behaviour. The stopping rule is fixed in advance precisely so nobody can stop at a flattering moment.
The usual causes: stopped early on a false positive, a primary metric that does not connect to revenue, a novelty effect that faded, a segment result generalised to everyone, or a win that was real but too small to see against normal variance. The design review usually finds it quickly.
Adobe Target makes sense if you are already invested in Adobe Experience Cloud. Optimizely is strong for server-side and feature-flag experimentation. VWO is often the best value for client-side programmes at mid-market scale. GrowthBook is worth considering where you want open-source and warehouse-native. The right choice depends on where you want testing to happen, not on the feature grid.
Only if you have enough behavioural signal to segment meaningfully and enough content to differentiate. Personalization on thin data mostly adds complexity. We size that honestly before recommending it, and we hold personalization to the same measurement standard as a test.
Yes — commonly on the statistics, the technical build or the measurement layer, which is where most in-house programmes are stretched thinnest.
Ready To Make Your Data Work Harder?
Let’s build a trusted measurement foundation that drives smarter decisions and measurable growth.