The client was hesitant about A/B testing, and wholesale orders running through its retail product pages had baked outliers into every checkout metric.
Here’s how Anatta de-risked the test, rebuilt a clean signal from raw order-level data, and delivered a ~2% AOV lift worth ~$1.27M a year – in a headless Shopify stack, with Convert Experiences.
Headline result:
+$2.12 AOV per order (~2%), with a 97.6% probability that the effect is real. At the brand’s order volume, this change projects to ~$1.27M in additional annual revenue – on existing traffic, with no additional acquisition spend.
The lift itself is modest. What stands out is where it came from: a test a reluctant client had to be coaxed into, run on a checkout where wholesale orders had baked in the outliers – and a result still clean enough to deploy.
Why read this case study:
Two challenges in this project will be familiar to almost anyone who runs experiments for a living.
The first is a client who isn’t yet sold on testing: the brand was hesitant about A/B testing and wanted any unproven experience in front of as little traffic as possible
The second is a dataset that won’t behave. Because the brand sold wholesale orders through the same product pages as retail, large and infrequent wholesale baskets had baked outliers straight into the checkout data, skewing the platform’s aggregate order value past the point of usefulness.
Anatta’s real work began with earning the client’s confidence in a test they would agree to run, and finding a way to read a clean result through contaminated data.
Where Anatta started:
The client is a direct-to-consumer brand whose best economics come from a bundle offer (“buy 3 for the price of 2”).
Paid traffic was being sent to a standard product detail page (PDP) that presented a single product variant first, leaving the brand’s most profitable option underexposed at the moment of highest intent. A large share of that traffic was mobile, where a cluttered or buried offer costs conversions disproportionately.
The hypothesis:
We believe that consolidating the PDP into a single tabbed buy-box that elevates the “Best Value” bundle, optimized for mobile, will increase AOV (and CVR) for high-intent shoppers, because the bundle has historically produced higher AOV and conversion than the single-variant PDP.
Before designing the test, Anatta built a customer model from historical interaction and order data to establish what “normal” looked like — average conversion rate, order-value distribution, and customer types.
The model:
- Grounded the team’s expectations of AOV and behavior (which later centered the priors in the statistical analysis)
- Surfaced a behavioral insight that wasn’t part of the hypothesis at all: new customers were significantly more likely to take the bundle offer than returning customers. That finding didn’t change the test – but it later helped further explain why the variant won.
The test design:
Control: the existing PDP, leading with a single product/variant.
Variant: a universal, all-in-one PDP with a tabbed buy-box.
The first tab presented the single product. The second tab — labeled “Best Value” – opened a bundle builder letting shoppers assemble the “buy 3 for 2” offer in one place. The variant was explicitly mobile-optimized, since a large segment of paid traffic arrived on mobile and the previous experience hadn’t prioritized that flow.
Because this was a headless Shopify environment (a modern Remix/Hydrogen-style front end), Anatta built the test on a server-driven experimentation foundation rather than a client-side toggle – which is what kept the experience flicker-free and the data trustworthy end to end.
Statistical rigor:
| Sample size per variation | 34,730 control / 6,144 variant (~85/15) |
| Conversion rate | 1.9% |
| Posterior Probability | 97.6% posterior probability of a positive AOV effect |
| Test duration | 25 days |
| SRM check | Observed split 15.03% consistent with configured 85/15 |
| Statistical method | Bayesian hierarchical model, log-normal likelihood (PyMC, Hamiltonian Monte Carlo / NUTS) |
The data source. All numbers were computed from raw, order-level data exported from the Shopify backend – not from the platform’s in-platform aggregate. Since wholesale orders flowed through the same product pages, the aggregated number was drastically skewed and had to be discarded in favor of order-level data with wholesale and outliers (and orders under $20) filtered out.
The model. AOV is right-skewed – many moderate orders, a few large ones — so a normal model would misrepresent it. Anatta used a Bayesian hierarchical model with a log-normal likelihood, with priors centered on historically grounded log-means (from the customer model) and kept wide enough that the live test data dominated the posterior.
In simple terms: The model was built so a few big spenders did not distort the average, a small segment of extreme results wasn’t given undue credence, and Anatta could begin from what the store already knew (priors) — while still letting the actual test results pick the winner.
Results: annualized return of $1.27 M
The variant produced a ~$2.12 lift in AOV (~2%), and the model put the probability of a genuinely positive effect at 97.6%.
The raw observed means tell the same story from the other direction: control AOV of $102.63 rose to $104.74, a $2.11 lift.
The fact that the model’s estimate ($106.28 → $108.40, +$2.12) and the raw sample means ($102.63 → $104.74, +$2.11) agree on the lift is the point – the log-normal model is stabilizing the estimate.
| Control | Variant | Improvement | |
| Posterior mean AOV (model) | $106.28 | $108.40 | +$2.12 |
| Observed AOV (raw backend) | $102.63 | $104.74 | +$2.11 |
| Probability of a positive effect | — | — | 97.6% |
| 94% credible range of the lift | — | — | $0.03 to $4.25 |
Contextualized to scale:
| Per-order lift | 25-day test volume (40,874 orders) | Annualized | |
| Expected | $2.12 | ~$86,653 | ~$1,265,000 |
Caveats:
1. First, the credible range is wide – $0.03 to $4.25. The lower bound sits just above zero, meaning there’s a small chance the real effect is minimal; by the same token, there’s an equal chance it’s meaningfully larger than $2.12.
2. The $1.27M figure is built entirely on the AOV lift, holding order volume constant – it assumes the conversion rate doesn’t change.
Segment-level findings:
The clearest segment signal came from the customer model and was validated against the test: new customers were significantly more likely to take the bundle offer than returning customers.
The variant didn’t just make a page prettier, it removed friction in front of the audience most receptive to the offer, which is why a structural PDP change moved a revenue metric.
Anatta also ran an additional analysis into Core Product Bundle orders specifically, to confirm the lift correlated with engagement on the bundle components rather than being an artifact of the smaller variant sample.
The surprising take-away:
The “safe” 85/15 split was the riskiest thing about the test, statistically. The client’s instinct – send only 15% of traffic to the unproven variant – was sound business risk management.
But a smaller arm carries more variance, which inflates error rates and makes true directional effects harder to detect.
The cost of the small variant group also shows up in the result. The fewer orders you send to a variant, the noisier its reading – and the harder it is to be sure a real gain is real.
This noise stretched the range of likely outcomes all the way down to $0.03 per order. A traditional significance test, with only 15% of traffic on the variant, would most likely have come back “borderline at best” The Bayesian approach Anatta used was built for exactly this: it could still pull a trustworthy, deployable result from a deliberately small arm.
Test-specific additional learnings:
In a mixed wholesale/retail Shopify store, the platform’s aggregate AOV is the wrong number. Wholesale baskets skew it badly. Order-level backend data – with wholesale, outliers, and sub-$20 orders filtered – is the only defensible basis.
This is also why the team reported AOV rather than RPV: RPV became unreliable once wholesale revenue had to be stripped out, while AOV cleanly answered the question — how much does this change move people and across which segments?
An AOV-only projection has a CVR assumption baked in. A revenue forecast built from a per-order lift × order volume holds the order rate fixed. State that assumption explicitly so a reader (or a CFO) knows the projection’s boundary.
The role of Convert Experiences:
Most A/B test write-ups treat the platform as a logo at the bottom. Here Convert Experiences was a partner that made the analysis possible.
Anatta implemented Convert’s Full-Stack SDK as a server-driven experimentation backbone. Instead of a client-side toggle, the server made the variant decision at render time and every downstream step inherited the same decision, keeping the data clean.
The capability that the Anatta team relied on the most was order-level experiment attribution.
Convert Experiences knew the visitor’s bucket, but advanced business analysis needed that experiment identity to reach the commerce backend.
Anatta propagated experiment context into cart attributes that flowed naturally into Shopify order metadata – so every order carried the variant that produced it. From this, the team computed observed AOV based on raw order-level data (rather than the skewed in-platform aggregate), filtered wholesale and outliers, and answered which variation drove higher-quality checkouts.
Around that core, Convert Experiences supported the production-grade controls:
- Single custom Goal API unifying storefront and checkout events (with a hardened CORS/validation model so Shopify Pixel could post trusted payloads)
- QA forcing flow to preview variants without polluting real traffic
- Bot and health-check suppression to prevent inflated views
- Stale-context protection so old experiment values didn’t linger across environment changes.
End-to-End run-time architecture:

What happened after the test?
After validation and rollout confidence, the winning variant became the default experience and the test-specific switching paths were retired – a transition completed in late 2025 as part of the cleanup toward defaulting to the winning V2 path.
Test was worked on by:
Anatta builds high-performance eCommerce systems where speed, conversion, and lifetime value matter more than aesthetics alone. They have purposely pivoted from the billable hours agency model, instead focusing on outcomes delivered according to business priorities and a collaboration experience that is as seamless as being an extension of the in-house team. Anatta brings a rare maturity to the e-commerce space, asking tough questions and improving what comes after the first sale.
Convert Experiences is the preferred A/B testing & personalization tool for lean growth & marketing teams, CRO agencies, and brands in tightly-regulated industries like insurance. GDPR-compliant, CCPA-ready, ISO 27001 and SOC-2 certified, it offers easy consent management, only first-party cookies, in-app warnings to maintain privacy, anonymization, data segregation and more at publicly-listed, affordable prices.


