Why Your Pricing Tests Keep Failing
Work the pyramid from the bottom up
When you think of first running price tests, what comes to mind? Adjusting your prices up or down? Maybe testing discount levels?
That’s what comes to mind for most of us, but it should be the last thing you test. Most subscription apps and SaaS companies test pricing in the wrong order, tuning the number while the foundations underneath it are broken. Then the test comes back flat, and everyone concludes that pricing isn’t the problem. Or they eventually change their model, and all the previous learnings go to waste.
Trust me, this isn’t my first rodeo.
After auditing and experimenting across subscription apps, I’ve landed on a simple mental model: pricing decisions stack like a pyramid. There are four layers, each dependent on the one below. Test them out of order and your results can mislead you, or you’ll find yourself having to rerun experiments.
Layer 1: The value metric (the foundation)
Before you touch price points, ask the harder question: what are you actually charging for?
Your value metric is the unit that scales your price:
- Hours of usage
- Number of seats
- Projects
- Messages
- Exports
It sounds like a packaging detail, but it’s not. It is the single biggest pricing decision you will make, because it determines whether paying more ever feels fair.
Here’s the test: does your value metric grow with the value the customer receives, or does it punish them for using the product?
A lot of apps price per usage, and with AI carrying compute costs, it feels smart. And on the surface it often looks like it works: heavy users hit the cap, upgrades get triggered, top-up purchases grow. The dashboards suggest everything is fine.
But the qualitative research often says something else entirely. Users describe feeling nickel-and-dimed. They tell you about the workarounds they’ve built to avoid hitting the cap. The value metric is generating revenue by creating the exact moment of frustration most likely to drive churn.
No price test fixes that. You need to get the fundamentals of your model right first, so payment aligns with value. That is where your tests need to start:
- What are you charging for?
- What is your pricing model?
Layer 2: Packaging (the structure)
Once you’re confident in what you charge for, you can look at how your plans are structured. This is where complexity creeps in and quietly kills conversion.
An app I once worked with had four tiers, each available monthly and annually. That’s eight options at the moment of highest drop-off. Users comparing plans were trying to do the maths and work out what was right for them. It was too much.
Look at Cloudflare Workers’ pricing page: 4 toggles, 4 options and 20+ features named… it’s too much.
Every extra option at the paywall or on your pricing page is cognitive load. Every inconsistency is a small leak of trust. I’m not saying don’t give options; rather, make sure every option earns its place. And if choosing the right plan requires users to predict their own future usage, you’ve moved the hardest question in the purchase decision onto the customer rather than guiding them.
At this level, you want to test things like:
- The number of tiers and what each one includes
- Plan length options (e.g. weekly, monthly, quarterly, annual, lifetime)
Layer 3: Price point (willingness to pay)
Now, and only now, does the actual number matter. And even here, most teams guess when they should measure.
Run proper willingness-to-pay research before testing price changes. Van Westendorp surveys (asks four questions to work out a range they are willing to pay) and MaxDiff analysis (having people choose the best and worst option) are cheap relative to what they tell you.
They also help you work out whether pricing is actually a constraint, and how much work the brand needs to do to close the gap between your desired price and what users are willing to pay today.
This is where you get to test around price, but please keep your tests clean. I’ve seen combination tests too often: a price adjustment and a change to the trial in the same experiment. When that test wins (or loses), you have no idea which change did it, so you can’t build on the learning.
Layer 4: Optimization (the top of the pyramid)
The top layer is where teams like to start: paywall tactics, anchoring, urgency, discount framing. These are legitimate levers, but they are refinements, not foundations. You can get creative here, testing different paywalls and what to include on your paywall vs changing your actual price.
They will build on your strategy and support it. But the major revenue lifts come from further down.
The order matters more than the tactics
The pyramid gives you a diagnostic sequence. When conversion or revenue underperforms, work upwards:
- Does the value metric reward or punish usage?
- Is the packaging simple enough to choose from in under a minute?
- Does the price match measured willingness to pay?
- Only then: is the presentation doing the price justice?
Layers one and two are where pricing problems live. Layers three and four are where pricing tests happen. That gap is why so many teams run test after test and conclude that their users are just price-sensitive.
They usually aren’t. The pyramid is just upside down.
Written By
Daphne Tideman
Growth Marketing, Direct-to-Consumer (D2C) Growth, Subscription App Optimization +4 more
Growth Marketing, Direct-to-Consumer (D2C) Growth, Subscription App Optimization, Conversion Rate Optimization, User Activation, Paid Social A/B Testing, User Research Recruitment Show less
Edited By
Carmen Apostu





