Testing with Feature Flags and Rollouts in Product Experimentation

Uwemedimo Usa
By
Updated ·

Feature flags and rollouts empower product teams to release features gradually, test in production, and drive data-driven product optimization.

Key points:

  • Feature flags turn features on or off, while rollouts use feature flags to gradually increase users’ exposure to new features.
  • Feature flags have moved from a software development technique into a core product experimentation tool, giving teams realistic insight into how a change performs with live traffic.
  • Practical applications and benefits of feature flags and rollouts include feature testing, personalized user experiences, and kill switches.

What are the chances a new feature release goes horribly wrong?

Higher than many teams are willing to admit. Features that pass every staging test still fail in production, where real traffic, real data, and real devices expose code paths no test suite covers.

So the safest releases control who sees a new feature once it is live. Feature flags make that possible by turning functionality on or off at runtime without shipping new code, and rollouts use those flags to widen access one subset of users at a time.

This guide covers what each one does, when to reach for which, how to test software once flags multiply the states it can be in, and how both work inside Convert Experiences.

What Are Feature Flags?

Feature flags are a technique that lets teams turn functionality on or off at runtime without deploying new code. They’re also known as feature toggles or switches. This allows for continuous experimentation and iterative development by controlling which users see the new feature and using data to decide whether to keep it.

Feature flag illustration

With feature flags, you can:

  • Reduce deployment risk. Deploy code without immediate release, enabling testing and quick rollback if needed.
  • Accelerate release cycles: by integrating hidden features into production early, enabling smaller and more frequent updates.
  • Conduct more realistic tests of new features with real users in production, for better data and faster feedback.
  • Run more targeted rollouts by controlling which users see new features, enabling strategies like canary releases and A/B testing.
  • Empower non-developers to control feature visibility. Product and marketing teams can toggle features and launch experiments without waiting on an engineering deployment.

There are nine main types of feature flags.

Types of Feature Flags

You’d find these pretty intuitive because feature flag types are named by what they do or how they behave:

Flag type Defined by
What it does
Typical use
Release Function Separates code deployment from feature release, deploy features into production without making them available to users Phased rollouts, canary releases
Experimentation Function Serves different versions to different user groups and measures impact A/B testing, personalization
Operational Function Adjusts system behavior in real time without touching the UI Incident response, maintenance mode
Permission Function Controls feature access by segment, role, or subscription tier Beta access, plan gating
Kill switch Function Shuts down a feature instantly without rolling back the release Emergency response to bugs or security issues
Long-lived Lifespan Stays in the system for extended periods Features under development or continuous monitoring
Short-lived Lifespan Bounded by a set timeframe, then removed from the codebase Single release cycle, quick experiments
User-based Scope Toggles by user attributes or behavior Personalization, dynamic pricing, role-based access
System-based Scope Toggles automatically on system conditions or events Load balancing, failover, dynamic security

At Albert, we used feature flags to run A/B tests on different user segments without disrupting the entire user base. This helped us gather real-world feedback and data on the new feature’s performance. For example, when we decided to revamp the user dashboard, we used feature flags to show the new version to just 10% of our users. Their interactions provided invaluable insights, which we used to make crucial adjustments before a wider rollout.

Will Yang, Head of Growth and Marketing, Instrumentl

Concrete Use Cases of Feature Flags

Feature flags have long been part of a developer’s toolkit. They evolved alongside practices like CI/CD (continuous integration and continuous delivery), where software can be released to production anytime.

As they became less of a niche tool, feature flags were integrated into comprehensive software delivery platforms, cementing their role in software development and product experimentation.

This has led to expanded use cases, such as:

1. Testing in production: Feature flags allow developers to test features directly in the production environment without impacting the entire codebase. This minimizes risks and provides a kill switch for when things go sideways.

This helps when you want to learn about the true performance of a feature in production, and staging environments just don’t cut it.

When we launched a new dashboard interface […], we used feature flags to limit its visibility to 5% of our users—those who regularly use several features every day. In addition to providing information about usability, this allowed us to determine how much the feature would increase the total system load. We might optimize the interface before a wider release by iteratively modifying it in response to direct user feedback and incremental performance statistics.

Moreover, feature flags support a strategy of gradual deployment. Following the first round of testing, we progressively increased the user base while modifying the functionality in response to ongoing input.

Alex Ginovski, Head of Product and Engineering, Enhancv

As two experts in the field note, understanding the feature’s true performance prioritizes both business and learning metrics.

In most cases, when we roll out the new feature, we are interested in both business performance and learning metrics.

The performance metrics highly depend on the feature and the type of business we deal with. Mostly, those are revenue-related metrics (Revenue per User, Revenue per Order, Sale Conversion rate, etc.).

Learning metrics mostly come from the micro-conversions that refer to certain interactions with the feature that we built. Quite often adoption rate is one of the focus metrics. Additionally, sometimes it is helpful to integrate NPS for the new features to learn about their acceptance by the target group.

Anastasia Shvedova, Lead Consultant Product Experimentation, and Michael Klöpping, Head of Engineering, KonversionsKRAFT

2. Canary releases: Roll out a new feature to a small subset of users before making it available to everyone else. This allows real-world testing and validation in a production environment with minimal fallout if things go wrong. Feature flags enable this. If necessary, you can quickly employ the next use case of feature flags.

During beta testing of a new AI-driven recommendation engine for our application, this was an effective application. We enabled the engine for a limited, targeted user group using feature signals while closely monitoring performance and user interaction. This enabled us to detect flaws and enhance the functionality prior to its wider implementation.

Jessica Shee, Senior Tech and Marketing Manager, M3 Data Recovery

3. Rollback/kill switch: There will come a time when you need a quick ‘undo’ button to remove a problematic feature post-release without impacting the entire codebase.

4. Server-side A/B testing: Toggle features for different user groups to compare performance. This is different from client-side A/B testing, where changes happen on the user’s web browser, for example. In fact, developers have used this to compare features’ performance long before cloud infrastructure and DevOps practices made it more common.

5. Feature gating: Facilitates targeting rollouts to specific user groups for beta testing, for example. It is also used for restricting feature access to certain users based on regulations, subscriptions, access levels, or user groups.

Other use cases include:

  • Minimizing disruption during updates
  • Feature sunsetting, and
  • Phased rollouts where you gradually introduce features to users in percentage increments

We integrated feature indicators with phased deployment for rollouts. The feature was initially enabled for 10% of users; this percentage increased incrementally as user confidence in its stability increased. A regulated approach was implemented to mitigate risk and guarantee a seamless user experience.

Jessica Shee

How Do You Test Feature Flags

Feature flags make releases safer, but they also multiply the states your software can be in

 One boolean flag gives you two paths to test. Ten flags give you 1,024 combinations. You’ll never test them all, so the job becomes working out which states actually carry risk.

(Note: The first three steps belong to whoever owns the test suite. The fourth is where the rest of the team comes in.)

Test Both Paths And Ignore The Flag

At the unit level, nothing changes. Write tests for the old behavior and the new one as separate functions, with no reference to the flag at all. If your flag picks between two search algorithms, test each algorithm on its own merits. The flag only determines which path runs, so your unit suite should pass regardless of which way it points.

Pick Three States At The Integration Level

This is where combinations get away from you, so choose the three that matter: what users see in production today, what happens if every flag falls back to its default (which is what you get when the flag service is unreachable), and the configuration you’re about to ship.

Add persona-based end-to-end tests only where a flag touches something you can’t afford to break, such as signup, checkout, or payment.

Let Production Tell You The Rest

Pre-release testing proves the code works. Production proves the feature works. Roll out to your own team first, then one percent, then widen while you watch error rates, latency, and the metrics you actually get judged on.

Before you go to everyone, run a non-inferiority test. This is an experiment that confirms the new feature hasn’t cost you anything and tells you by how much. This is where a feature flag stops being a QA tool and starts being a product experimentation tool, because the question changes from “does it work?” to “is it better?”

Delete The Flag When It Has Done Its Job

A flag that has been at 100 percent for a month is technical debt. Strip out the flag logic, delete the path nobody takes anymore, and archive the flag. Your test matrix shrinks back down with it.

How Do Feature Flags Work in Convert Experiences?

This video demonstrates how feature flags work within Convert’s Fullstack product:

Convert provides six Fullstack SDKs:

  • JavaScript/TypeScript
  • PHP
  • Python
  • Ruby
  • iOS
  • Android

All six share the same feature-flag, rollout, targeting, and bucketing model, so a visitor is bucketed the same way, whichever one you integrate.

What differs is the language-specific API and how each SDK handles persistence, configuration, and event delivery. The JavaScript SDK also covers edge environments such as Cloudflare Workers.

To use feature flags in Convert, start in SDK Config & Keys inside your Convert Fullstack project, which generates a code snippet for the SDK you’re using.

Public keys are safe for client-side JavaScript, but authenticated keys contain a secret and should only be used in server or edge environments. Each key is assigned to Live, Staging, or both, so staging traffic stays out of live reporting. From there, define the specific areas and features in your app to target for the experiment.

Each feature has a key and a set of typed variables (boolean, integer, string, float, or JSON), and the SDK returns both the feature’s status and its variable values to your application. You use those values to serve your specific variation customization logic. You can also toggle these features within the code.

Optionally, set up event listeners to monitor feature activation and user interactions for debugging.

Once set up, Convert assigns users to different feature variations using the user ID you pass in. The bucketing is randomized yet deterministic for a specific user context, so the same identifier returns the same variation as long as the configuration is unchanged. Consistent assignment across sessions depends on passing a stable user ID or configuring a DataStore. You can track conversions to analyze feature performance with preset and custom goals.

All of this data is accessible through Convert’s robust reporting and analysis tool.

What Are Feature Rollouts?

Feature rollout is a software development strategy for gradually introducing a new product or feature to users. Its main goal is to ensure the product—whether a minimum viable product (MVP) or a fully developed feature—performs as intended before reaching the entire user base.

This is closely linked to progressive delivery, which involves releasing new features to a smaller user group first to assess their impact before a wider release.

Feature rollout illustration

Feature rollouts control and phase the release of a validated feature, ensuring that the entire product system can handle the increased usage load.

Anastasia and Michael explain how they do this at KonversionsKRAFT:

Sometimes the step [after a successful feature A/B test and] before the native development would be to put the variant on 100% of audience traffic.

However, we highly recommend limiting the timeframe for this adjustment for two reasons:

  1. For some tools this means really high costs due to the increased amount of impressions.
  2. The experiment code is still fragile and the changes on the control-side can affect the variant functionality.

What we practice a lot is the phased rollout strategy that follows the following logic.

We might:

  • Rollout to a smaller amount of users and then gradually increase the amount of the affected traffic (soft roll out)
  • Rollout to one segment of users (for example if we want to reduce the quality assurance effort). The segmentation can be both technical (by browser/device) or socio-demographic (a customer group with a specific consumer behavior).
  • Rollout in one country as a starting point.”

Anastasia Shvedova and Michael Klöpping

This careful management helps maintain system stability and performance as you introduce new features. On top of that, you can also:

  • Catch bugs early, for faster fixes and a better user experience.
  • Minimize potential issues for the entire user base.
  • Enable real-world testing with actual users, providing valuable feedback for improvements.
  • Personalize user experience by tailoring feature rollouts to specific user segments.
  • Acquire data on how users interact with new features, a key asset for product optimization.

Note that while A/B testing is used to evaluate untested ideas and release changes to users, they typically involve a fixed 50-50 randomization of control and treatment groups. In contrast, rollouts offer more granular control of the user groups for a phased introduction of features to the user base.

How Are Rollouts Used?

Here lies the key difference between feature flags and rollouts.

Feature rollouts typically follow a successful experiment. They are used to ramp up traffic away from the control treatment whilst system and customer behaviour metrics are monitored. As an example, when re-platforming technology, an experiment between a new variant can de-risk the decision to move, and the roll-out will de-risk the actual move.

Andrew Brimble, Experimentation Lead at Specsavers

Think of it this way: Feature flags are like light switches that can turn a feature on or off. Rollouts, then, are like dimmer switches that use feature flags to gradually increase the brightness of a new feature so more and more users can experience it.

Makes sense?

That said, you’d use that ‘dimmer switch’ in the following cases:

  1. Targeted user feedback: Release a feature to specific user segments to gather targeted feedback from those most likely to use it. Use this input to refine the feature and better meet user needs.
  2. Personalization: Provide distinct user experiences to specific segments (based on preferences, behavior, or demographics).

We once conducted a successful experiment for a client in the field of educational e-commerce. User research indicated that different target groups on the website had varied needs regarding category preferences. To address this, the team proposed a feature that allowed personalized views of selected categories in user accounts.

Given the complexity of this new journey step and its integration with backend logic, implementing it as an A/B test was deemed too fragile. Instead, we rolled out the feature as a flag for 100% of selected user segments. This approach allowed us to control the feature’s availability dynamically.

The results were impressive: interaction rates with categories increased significantly across target groups. Consequently, our client decided to implement the feature fully in one country as a pilot.

Anastasia Shvedova and Michael Klöpping

  1. Controlled testing: Test with a limited audience before releasing to a large user base, and eventually the entire user base. This way, you spot bugs, usability problems, etc., without impacting everyone’s experience.
  2. Limited impact release: If a new feature has unforeseen negative consequences, a rollout limits the number of affected users, making it easier to roll back the change or implement fixes quickly.

At Instrumentl, we combined feature flags with gradual rollouts to mitigate risks and catch any issues early on. For example, when rolling out a new grant-tracking feature, we started with a feature flag to test it in a controlled environment.

Once the initial bugs were ironed out, we used a phased rollout strategy to incrementally expose the feature to more users. This approach not only protected user experience but also provided multiple checkpoints to optimize the feature. The key is to continuously monitor user feedback, making adjustments as needed, ensuring the feature enhances user satisfaction and meets business goals.

Will Yang

Keep in mind that incorrect usage of feature rollouts (and even feature flags) can be detrimental to your business and users.

For example, if there isn’t a clearly defined process for removing feature flags after they have served their purpose, such as after confirming the feature is widely accepted by users and stable in the production environment, you risk building up technical debt. These flags can clutter your codebase and make maintaining the code difficult.

If an A/B test was successful, we recommend our clients to proceed with the native development of the feature on their page.

If we did an experiment with the help of a feature flag, then this step is not required since [we developed] natively from the very beginning, the feature rollout becomes much easier. However, our recommendation is still to make a code cleanup and ensure that everything looks correct from the technical perspective.

Anastasia Shvedova and Michael Klöpping

Also, failing to monitor key performance indicators and health metrics during a rollout can mask underlying problems and negatively impact user experience, even when a feature appears successful in the short term.

We try to ensure a smooth feature/test rollout on several levels.

First of all, on a technical and quality assurance level. Once the test is developed we always proceed with the code review and unit test to ensure the functionality of our code.

The next thing that we do is the so-called exploratory testing with native devices. In our opinion, nothing can replace testing by a human when it comes to identifying some UX issues. Once the test is “tested” and launched, we do a go-live monitoring.

Usually, it includes steps like checking the code environment, checking the console, and making sure that all the metrics are firing correctly. We also create the “Error goals” as metrics based on the “try-catch” method — it helps us identify that the elements from the test are displayed and function correctly.

Secondly, on the metrics level. Of course, we look at the business metrics that help us see how our feature or test affects the business performance. Additionally, we often try to roll out a non-inferiority test experiment before we launch the new feature. It helps us answer the question of whether the new feature is actually not ruining some business metrics and, if yes – to what extent.

Anastasia Shvedova and Michael Klöpping

How Do Rollouts Work in Convert Experiences?

This video shows how Convert’s Fullstack Rollout feature enables controlled feature releases through

  1. Easy activation of features by changing rollout status
  2. Immediate release of a feature to all users, once your application code is already deployed
  3. Simple and controlled feature management, allowing for gradual rollouts adjusting traffic allocation in the dashboard.

A FullStack project in Convert is built so engineers and marketers work on the same rollout from opposite ends. Engineers integrate one of six SDKs (JavaScript/TypeScript, PHP, Python, Ruby, iOS, or Android), define features and their typed variables in Convert Experiences, then read them in code, and use a stable user ID or a configured DataStore so bucketing remains consistent for a given user across sessions.

Marketers and product teams then control the release itself from the Convert dashboard. They can switch a rollout on or off, adjust the percentage of traffic exposed, and read results against preset or custom goals, without filing an engineering ticket or waiting for a deploy. The code ships once, and the decision about who sees it stays adjustable afterward, which puts the people closest to the customer in charge of the release.

Feature Flags vs Feature Testing

Feature flags and feature testing both support software development, but they’re not the same thing.

Feature flags are conditional logic statements embedded in the code to switch specific functionalities on and off during runtime, without redeploying code.

Feature testing encompasses the methods and processes used in evaluating the performance, stability, and user acceptance of new features—before or during rollout.

While feature flags can be used to ‘test’ the impact of having vs not having a feature (on vs off), feature testing goes wider, enabling experimenters to vary certain feature parameters, one at a time, to understand user preferences and monitor the impact of different configurations on predetermined KPIs.

To tie it all together, Andrew Brimble describes how Specsavers use feature flags, rollouts, and feature testing to optimize their single page application:

Specsavers utilizes feature flagging within their in-house booking app, where customers book sight and hearing appointments at one of our many high street stores. This highly optimized single page app is responsible for over 80% of booked appointments. Its value to the business commands a rigorous approach to experimentation. For non-trivial optimizations or new feature roll-outs, Specsavers have chosen to integrate the experimentation within the product engineering teams. Oftentimes, client-side testing cannot achieve what is needed on the back end to achieve the desired change. The 3 main benefits have been:

1. Minimise code error: No conflicting code injected into the browser, follows standard development process, including extensive testing and release management.

2. Experiments can run in multiple markets: Our global codebase allows the release of experiments worldwide.

3. ‘Winning’ variants can be rolled out with ease: Usually the removal of the feature status logic is all that is needed.

With reference to point 3 above, you may decide that front-loading development of an un-tested feature is too costly. When a minimum viable product is tested more effort is needed following the experiment. This problem is not exclusive to feature flag testing, and the MVP code may not always be thrown away entirely.

Andrew Brimble

In a Nutshell

Feature flags and rollouts may present as just safety nets, but in the hands of the right product team, they’re true catalysts for product-led growth. They enable faster, safer releases, targeted user feedback, and data-driven optimizations that ensure you ship your most refined features.

Master these tools and the different roles they play to build products that truly resonate with your users.

FAQs: What Are Feature Flags and Rollouts and How Are They Used in Product Experimentation

1. How do you integrate feature flags with A/B testing?

Feature flags and A/B tests answer connected questions. A flag decides which code path a user receives, and an A/B test measures which path performs better. Integrating them means passing the same stable user identifier to both the flag evaluation and the conversion tracking, so exposure and outcome are attributed to the same person. A FullStack project in Convert Experiences handles both from a single SDK integration, resolving the feature decision and recording conversions against preset or custom goals in the same project.

2. What is a non-inferiority test in a feature rollout?

A non-inferiority test is an experiment designed to confirm that a new feature does not degrade the metrics a business already relies on, and to quantify the extent of any degradation. Teams run one before expanding a rollout beyond a small traffic percentage, when the goal is to prove safety rather than improvement. The test defines an acceptable loss margin in advance, often a fraction of a percent of revenue or conversion rate, and the rollout proceeds only if results stay within it.

3. Do feature flags slow down your application?

Feature flags add negligible latency when the SDK holds its configuration in memory and evaluates decisions in-process, but latency becomes measurable when every check requires a network round trip. SDK placement matters more than flag count. Convert Experiences supports edge evaluation via Cloudflare Workers, keeping the decision close to the request rather than routing it back to a central service.

4. How long should a feature flag stay in your codebase?

A release flag should be removed once it has been at 100 percent for a full release cycle, typically two to four weeks after full rollout. Flags left in place past that point become technical debt. That is, they widen the test matrix, keep dead code paths alive, and make later debugging harder. Permission flags and operational flags are the exception, since they are designed to stay in the system indefinitely as part of how the product works.

5. Do feature flags deploy code to your application?

Feature flags do not deploy code. Your application must already contain the implementation and read the SDK result; the flag decides whether that code path is active and which configuration a given user receives. In Convert Experiences, a Feature controls the status and variable values returned to your application, and it does not change your interface on its own.

6. What programming languages does Convert Experiences support for feature flags?

Convert Experiences provides six Fullstack SDKs: JavaScript/TypeScript, PHP, Python, Ruby, iOS, and Android. Each has its own quickstart in Convert’s developer documentation, and the JavaScript SDK also covers edge environments such as Cloudflare Workers. Feature flags in Convert Experiences run inside a Fullstack project, which is separate from a Web Testing project and its browser tracking script.

7. How do you keep the same user in the same variation across sessions?

Consistent assignment depends on passing a stable user ID, or configuring a DataStore that persists the assignment. Convert Experiences buckets users deterministically for a given user context, so the same identifier produces the same variation as long as the configuration is unchanged. Anonymous visitors whose identifier changes between sessions can be assigned to a different variation, which is why identity strategy belongs in the planning stage of a rollout.

8. Should you use separate feature flag configurations for staging and production?

Separate environments keep experimental configuration and test traffic out of production reporting. Convert Experiences supports this through SDK keys assigned to Live, Staging, or both, ensuring that a staging integration cannot write results to a live project’s data. Convert Experiences also distinguishes public SDK keys, which are safe for client-side use, from authenticated keys that carry a secret and belong only in server or edge environments.

CTA Full stack
Mobile reading? Scan this QR code and take this blog with you, wherever you go.
Updated - Originally published
Written By
Uwemedimo Usa
Uwemedimo Usa
Uwemedimo Usa
Conversion copywriter helping B2B SaaS companies grow.
Areas of expertise
Conversion Rate Optimization, Conversion Copywriting, B2B SaaS Content Marketing
+3 more
Edited By
Carmen Apostu
Carmen Apostu
Carmen Apostu
Content strategist and growth lead. 1M+ words edited and counting.
Start your 15-day free trial now.
  • No credit card needed
  • Access to premium features
You can always change your preferences later.
You're Almost Done.
What Job(s) Do You Do at Work? * (Choose Up to 2 Options):
Convert is committed to protecting your privacy.

Important. Please Read.

  • Check your inbox for the password to Convert’s trial account.
  • Log in using the link provided in that email.
  • To ensure you receive your 30-day trial from our ambassador, please use the same browser to claim your account.

This sign up flow is built for maximum security. You’re worth it!