Automated A/B testing: the AI finds the test and opens the PR
Friction found in your own session replays becomes a hypothesis, a pull request with variant B, and a split test that only names a winner once the statistics allow it.
How it works
- 1
Friction is detected
The same replays and heatmaps you already record are scanned every day for eight named failure modes: dead clicks, rage clicks, a buried CTA, deep scroll with no conversion, message mismatch, form abandonment, a mobile gap and attention leaks.
- 2
A hypothesis and a success goal
Each finding becomes a proposed test ranked by impact, confidence and ease, with the evidence rows that triggered it. You edit the hypothesis, pick the goal that counts as success, and confirm.
- 3
The AI opens a pull request
A background run reads your connected GitHub repo, implements variant B as a CSS-only change gated on the visitor’s arm, and opens a PR. It never commits to your default branch; your review is the approval.
- 4
Split, measure, wait
Visitors are assigned per person before first paint, so there is no flicker. The results page hides the confidence column until the minimum sample and 14 days have both passed, then names the winner.
What you get
The loop from a rage click to a merged fix, with the guard rails a statistician would insist on.
Eight failure modes, ranked by ICE
The CRO expertise is deterministic, not a prompt: each mode has a behavioural fingerprint and a real sample floor, so ten QA clicks can never read as "45% rage". Findings are ranked by impact, confidence and ease.
Flicker-free per-visitor assignment
A tiny inline bootstrap stamps the arm on the page before first paint from the visitor id, so the same person always sees the same variant and nothing flashes. The server re-derives the arm, so a tampered client cannot stuff a result.
Before/after when a split is impossible
A full redesign, a pricing change or a new SEO title cannot be shown two ways. Those run as a before/after with matched, equal-length windows, and the result is labelled confounded wherever it is shown.
Real statistics, no peeking
A two-proportion z-test with lift confidence intervals and sample-size planning up front. A winner is never named on p-value alone: minimum sample and two full weekly cycles are both required.
Results land in Insights
The proposed test, the running experiment and the verdict all appear as cards in the Insights feed, with a deep link into the analyst if you want to stop, ship or revert with a sentence.
Needs GitHub connected
The variant is real code in your repo. Connect GitHub, pick the repository and framework once, and the AI works from what is actually there. No visual editor, no injected snippet that a redeploy wipes out.
The eight failure modes it looks for
- Dead clicks
- Rage clicks
- Buried CTA
- Unconvinced (deep scroll, no conversion)
- Message mismatch
- Form abandon
- Mobile gap
- Attention leak
Every rule needs at least 50 clicks and 100 pageviews on the page, and a dead-click suspect needs eight distinct visitors, before it can propose anything.
Where VWO or Optimizely is the better choice
A capability page you can trust has to cut both ways. Honestly:
- There is no visual or WYSIWYG editor. VWO, Optimizely, AB Tasty and Convert all let a marketer change a headline by clicking on it. Here the variant is a pull request, which is what makes it survive a redeploy, and also what makes it a developer-adjacent workflow.
- You need a codebase on GitHub and someone to review the PR. If your site lives in a page builder with no repository, or nobody on the team merges pull requests, a hosted editor is the better tool for you.
- No multivariate tests, no personalization engine, no feature flags and no server-side SDKs for apps. Optimizely and VWO cover those; we run A/B splits and before/after measurements on websites, and nothing else.
- It is a Scale-plan feature at $199/mo, and it needs the session recordings that start on Growth. VWO has a free tier and Convert starts under $100. If you are testing one landing page a quarter, the cheaper editors are the right size.
The A/B testing tools, side by side
Split-testing tools compared on price, who writes the variant, and what it takes to call a winner.
| Real-world price | Visual editor | Who builds the variant | Finds the test for you | Session replay included | |
|---|---|---|---|---|---|
| BusinessMCP | $199/mo (Scale) | No | AI opens a GitHub PR | Yes | Yes |
| VWO | Free Starter · paid ~$314/mo+ (quoted) | Yes | You | No | Add-on |
| Optimizely | Custom (~$40k–$150k+/yr) | Yes | You | No | No |
| AB Tasty | Custom (~$15k–$150k/yr) | Yes | You | No | No |
| Convert Experiences | $299/mo annual ($399 monthly) | Yes | You | No | No |
Full teardowns
Frequently asked questions
What does "automated" actually mean here?
Three things are automated that are usually manual: finding the test (a deterministic engine scans replays and heatmaps for eight known failure modes), building the variant (the AI implements variant B and opens a GitHub pull request), and calling the winner (statistics with a peeking guard). Two things are deliberately not: you confirm the hypothesis and goal before anything is built, and you review and merge the PR. The AI never commits to your default branch.
Will visitors see a flash of the original before the variant loads?
No. Assignment happens in a small inline script that runs before the page paints, using the visitor id already in local storage, and stamps the arm onto the document. The variant is CSS gated on that attribute, so the browser never renders the control first. The tracker later re-derives the same arm from the same id, and the server does it again when it records the exposure.
When is a winner declared?
Only after both conditions hold: the planned minimum sample has been reached, and the test has run at least 14 days so two full weekly cycles are covered. Until then the results table hides the confidence column entirely. A test with a p-value under 0.05 on day three is a test you were about to be fooled by.
What if my change cannot be split, like a pricing change or a redesign?
Choose the before/after method. It compares equal-length windows on either side of the change, and every result is labelled confounded, because seasonality and traffic mix can move the number as much as the change did. Split traffic is the default; before/after is the honest fallback, not a hidden downgrade.
How much does it cost and what do I need?
A/B testing is included in the Scale plan at $199/mo, which also carries 30-day session recordings, the AI SDR and 10k monthly visitors. Prerequisites: the tracking script installed, heatmaps and recordings switched on, a success goal defined, and GitHub connected with the repository chosen. The Free plan lets you explore the rest of the product; the CRO engine itself starts on Scale.
Companies running MCP servers with us
Makers and companies publishing a server in the BusinessMCP directory.
Grabbit MCP
grabbit.sh
Find Reddit buyer-intent threads, inspect subreddit rules, read full…
FinancialData.Net MCP Server
financialdata.net
Get real-time stock prices, fundamentals, institutional trading insights, and…
API2Cart MCP
api2cart.com
Unified MCP server for 70+ eCommerce platforms.
Polyrama
github.com/Polyrama
Prediction-market terminal and MCP for live odds, market search, trader analytics…
Twitter/X API by SocialData
socialdata.tools
Read-only Twitter/X data for AI agents: search tweets, look up profiles…
MachineTranslation.com MCP
machinetranslation.com
Agents can translate text through 22 leading AI models at once and get back only…
PolymarketScan
polymarketscan.org
Free no-auth Agent API plus an independent record of Polymarket wallets and prints.
Humanizer PRO MCP
texthumanizer.pro
Hosted OAuth MCP for rewriting authorized draft text, scanning AI-style signals…
Process Street MCP Server
process.st
Connect AI agents to Process Street workflows, tasks, runs, data sets, and…
Magic Hour MCP
github.com/magichourhq
Generate and edit video, images, and audio through 44 Magic Hour tools from…
Trendos
trendos.com
AI search visibility analytics and opportunities.
ContextStream
contextstream.io
Shared project memory for AI coding agents across Cursor, Claude Code, Codex…
Asyntai – AI Support Agent for Websites
github.com/asyntai
Agents run a website's Asyntai AI support agent: manage its knowledge base and…
iLook
ilook.fit
Privacy-first, non-medical AI face-analysis tools for consented photos, with clear…
HelpMyAgent
helpmyagent.com
Business data, intelligence and procurement APIs designed for AI agents.
FoxForm
foxform.app
Agents can build forms, quizzes and calculators where each answer scores points…
Google Ads MCP Server
get-ryze.ai
Solana Sniper Bot MCP
github.com/Solara-Snipe-Bot
Autonomous Solana meme-coin sniper with 48 AI MCP tools and multi-source launch…
padel.how Racket Reviews
padel.how
CorpusIQ
corpusiq.io
Business operations platform exposing 65+ MCP tools for data, commerce, marketing…
CareClinic Health Tracker
careclinic.io
Track symptoms, medications, and mood from your CareClinic account, then review…
Vibe Prospecting
vibeprospecting.explorium.ai
B2B company and contact intelligence for prospecting workflows.
Let the replays pick your next test
Install one cookieless script, switch on recordings, connect GitHub. The first proposed test arrives once a page has enough real clicks to trust.
Month-to-month. Every variant is a pull request you review.