> Markdown version of https://www.getunleash.io//blog/operational-kill-switches-production
> For clean Markdown of any blog post, append .md to its URL.
> For a site index, see https://www.getunleash.io//llms.txt.

# Feature flags for running production: kill switches and dynamic configuration

_Published 2026-10-08 by Alex Casalboni in Industry Insights._

Most incident runbooks have the same step near the top: roll back the last deploy. It assumes the trouble arrived with that deployment, and that a fresh build can reach production before customers notice.

Neither assumption held for [Mintlify](https://www.getunleash.io/case-study-mintlify) on May 20. At 12:38am Pacific, engineering manager Nick Khami got paged because every MongoDB replica had gone down. The trigger was a marketing leaderboard page that queued a heavy refresh job on every load, and repeated visits stacked up more than 32,000 aggregation queries until the database gave out. The team got the site back with a code fix force-merged at 2am and containers restarted by hand, after roughly three hours of downtime spread across two outages. In Khami’s words:

> “Our only controls that night were git push and the AWS console.”

When the deploy pipeline doubles as your rollback mechanism and your incident response, every fix waits for a build and a review. The examples below are kill switches and runtime settings that let you change how production behaves within seconds.

You’ll find another post dedicated to financial services use cases in [operational feature flags for financial services](https://www.getunleash.io/blog/operational-feature-flags-financial-services).

## Kill switches for dependencies

Every external call on a critical path should have its own switch, with a pre-approved fallback.

### Checkout dependencies in e-commerce

Before a shopper reaches the pay button, a typical checkout has called a tax engine, a shipping rate API, an address lookup, a fraud screen, a buy-now-pay-later provider and a reviews widget. Each of those vendors has maintenance windows and bad days. A slow response from any one of them can hold the whole page hostage.

Put each call behind its own kill switch and agree on the fallback ahead of time. Nobody misses the reviews widget for an hour. When the carrier API times out, a flat-rate shipping table keeps orders flowing, and BNPL can quietly drop out of the payment options while cards keep working.

Tax is the delicate one: charging the wrong amount has legal consequences, so a cached rate per region plus a marker for later reconciliation is the safer path.

Store the fallback in a [variant payload](https://docs.getunleash.io/concepts/strategy-variants#variant-payload) on Unleash and you can change it mid-incident. If the carrier outage drags on into the evening, switch the fallback from flat rate to free shipping over $50 without shipping code.

### SMS, email and maps

One-time passcodes and delivery routing both depend on utility APIs that companies rarely think about until they fail. A food delivery app that loses its geocoding provider starts sending couriers to the wrong side of town within minutes.

A kill switch per provider lets you route traffic to a secondary vendor you’ve already contracted and tested. With no backup vendor, the switch can change the product instead: ask the customer to drop a pin on the map, or send the passcode by email while SMS is down. In Unleash, constraints on the country field will keep the fallback limited to the markets where the primary provider is failing.

### Identity providers and social login

Any fallback that touches login gets designed and approved with your security team. The kill switch picks between options that already meet your security bar, so a degraded login path is still a safe one. Think of it as graceful degradation for authentication, planned in daylight.

For a consumer app, the common case is a social login outage. If Sign in with Apple or Google stops responding, the switch hides that button and points people to the email login or magic link they’ve already set up.

Workforce apps behind a corporate identity provider face a sharper version of the problem, since an outage can lock out the engineers who need to fix things. Some teams keep a break-glass sign-in path for a short list of named on-call admins, gated by a [change request](https://docs.getunleash.io/concepts/change-requests) and recorded in the [audit log](https://docs.getunleash.io/concepts/events).

The same switch helps during a security incident at the provider. If your identity vendor reports a compromise, you can stop accepting its tokens across every service at once and move users to the backup method while you investigate.

### LLM providers

AI features add a dependency with its own rate limits and outages. Quality can also shift overnight when the provider updates a model, and a travel site’s trip planner or a telecom support assistant will show it to customers first.

Give each AI feature a kill switch with a variant payload naming the fallback. One option routes requests to a second provider or a smaller self-hosted model. Another turns the feature off with an honest message and sends support conversations to the human queue, so the rest of the product keeps working.

A switch can also flip itself. The Unleash team [runs its internal Slack agent](https://www.getunleash.io/blog/controlling-an-ai-agent-at-runtime-using-unleash) behind a kill switch with a [safeguard](https://docs.getunleash.io/guides/getting-started-release-management) attached, which shuts the agent off when its costs over three hours reach roughly ten times a normal period.

### Integrations and internal services

Some of the riskiest dependencies live inside your own company. After the May outage, [Mintlify extended flags to its back end](https://www.getunleash.io/case-study-mintlify). Nine days later an engineer shipped a switch that pauses the documentation update pipeline so jobs wait safely in a queue when the database is under strain.

Mintlify now runs eight standing kill switches, covering the caching layer, billing circuit breakers, indexing and the GitHub sync that keeps customer docs current. In a retail app, the equivalent targets might be the loyalty points service or the recommendations engine. If either one slows down, the switch hides the points balance or the carousel and the account page still loads.

Monoliths benefit too. [Tink, A Visa Solution](https://www.getunleash.io/visa-feature-flag-case-study) runs a large monolithic platform, and flags let its teams switch off one misbehaving feature without rolling back a deployment that carries everyone else’s work.

## Kill switches that stop money leaking

Some incidents leave every dashboard green and cost money every minute they run. The people who spot them first tend to sit outside engineering, so the switch belongs within their reach.

### Promotions and pricing mistakes

A discount code meant for 500 newsletter subscribers ends up on a coupon site at 9pm on a Friday. Or a data entry slip lists a television at a tenth of its price, and by the time anyone notices, screenshots are all over social media.

Put every campaign behind its own flag and give the merchandising team a [role](https://docs.getunleash.io/concepts/rbac) that lets them switch promotions off and touch nothing else. In Unleash, constraints on channel or market fields let them shut down the abuse in one place while the campaign keeps running for the newsletter list.

### Sportsbook markets

Live betting depends on an odds feed that keeps pace with the match. When the feed lags, bettors watching a faster stream can place bets on outcomes they’ve already seen.

Traders need to suspend a single market, or in-play betting across a whole jurisdiction, within seconds. Constraints on event ID and jurisdiction give them that precision. When a gambling regulator later asks exactly when a market was suspended, the audit log has the timestamp and the name.

### Exploits in live games

Players discover an item duplication glitch on a Saturday, and by Sunday the in-game economy is flooded with counterfeit swords. Patching the game client means a store submission on several platforms, each with its own review queue.

A server-side kill switch closes the player marketplace or disables one crafting recipe while the economy team investigates. Constraints on game version or server region keep the rest of the world playing normally.

### Billing circuit breakers

Usage-based billing jobs fail in expensive ways, like charging a customer twice or invoicing from a broken usage pipeline. For example, [Mintlify](https://www.getunleash.io/case-study-mintlify) keeps billing circuit breakers among its standing switches for that reason. When the numbers look wrong, the switch stops the charge job and holds invoices for review, and finance gets a look before any money moves.

## Kill switches for automation

Automation acts faster than anyone can review it, so every automated actor needs an off switch that works without a deploy.

### Repricing engines

Marketplace sellers run repricers that follow competitors up and down all day. A bad data feed, or two repricers chasing each other, can push a product down to one cent overnight.

A switch per category or seller freezes prices at the last value someone approved. Once the feed is fixed, the pricing team restarts the engine one category at a time with a [gradual rollout](https://docs.getunleash.io/guides/gradual-rollout) and watches how it behaves before widening.

### Airline rebooking

During a snowstorm, automated rebooking moves thousands of passengers onto new flights within minutes. If a bug starts placing people on flights that are themselves cancelled, every extra minute of automation creates more work for gate agents.

Stopping the engine for one hub, or for one day of departures, hands those passengers to the agents while the rest of the network keeps rebooking automatically. In Unleash, with constraints on airport and departure date, that’s a single change with no redeploy.

### AI agents with write access

Agents that issue refunds or change bookings need the same treatment. A switch per tool lets you take away one capability, such as refunds, while the agent keeps answering questions. We went deeper on this pattern in [securing AI agents starts with a kill switch you control](https://www.getunleash.io/blog/financial-service-institutions-ai-kill-switch).

## Kill switches where a redeploy takes days

Code running on someone else’s device can take days or weeks to replace, which leaves a server-side switch as the only fast lever.

### Mobile apps

An app store review can take a day, and plenty of users never turn on automatic updates. A broken feature in version 5.2 can stay in circulation for months after 5.3 fixes it.

Wrap new mobile features in flags evaluated against the app version and operating system, using the dedicated Unleash SDKs for [Android](https://docs.getunleash.io/sdks/android), [iOS](https://docs.getunleash.io/sdks/ios), [Flutter](https://docs.getunleash.io/sdks/flutter) or [React Native](https://docs.getunleash.io/sdks/react-native). When crashes spike on one Android build, you switch the feature off for that build alone and leave it running everywhere it works.

### Smart TVs, kiosks and warehouse devices

Smart TV apps, set-top boxes, airport check-in kiosks and handheld warehouse scanners all update on their own schedules. A hardware vendor or a store IT team often controls the timing, in waves you can’t speed up. TV and kiosk apps built on web technologies can use the [JavaScript SDK](https://docs.getunleash.io/sdks/javascript-browser), and our [SDK overview](https://docs.getunleash.io/sdks) covers the other platforms.

For example, [Mercadona](https://www.getunleash.io/case-study-feature-ops-mercadona) runs the logistics systems behind Spain’s largest supermarket chain, where a failing picking system can stall an entire shift of warehouse workers. Its logistics staff engineer describes recovery like this: “_I only have to switch off a feature toggle._” In Unleash, constraints on device model or site ID let you do the same for one warehouse or one generation of TVs.

### Zero-day vulnerabilities

When a critical vulnerability lands in a widely used library, as Log4Shell did in December 2021, the patch takes time to build and roll out across every service. During that window, the vulnerable code path is open to anyone scanning for it.

By wrapping risky input handling in flags, you can close those paths in seconds and reopen them after the patch lands. File uploads, document rendering, image processing and public webhooks are good candidates. Set each flag’s default in code so the risky path stays closed whenever the Unleash SDK has no flag state to work from.

## Peak events and recovery time

The busiest days of the year are known weeks in advance, so the switches you’ll pull on those days can be agreed in a calm meeting.

### Black Friday and ticket onsales

A retailer can decide the order in which features get switched off when traffic climbs. Personalized banners and live stock counters go first, followed by recommendations, search facets and finally anything rendered on demand. Each step buys capacity at a known cost to the experience. We covered the engineering patterns in [graceful degradation in practice](https://www.getunleash.io/blog/graceful-degradation-featureops-resilience).

A ticketing site facing a stadium onsale can put a virtual waiting room behind a flag and switch off seat-map previews for the first hour, when demand is at its highest.

The flag system has to survive the same peak. [Wayfair](https://www.getunleash.io/wayfair-feature-flag-case-study) serves more than 20,000 requests per second on a typical busy day through one Unleash server and a fleet of [Unleash Enterprise Edge](https://www.getunleash.io/unleash-enterprise-edge) instances acting as a distributed read-only proxy, and its SRE team reports that failures stay contained during Black Friday spikes.

### What changes once the switch exists

[Samsung Ads](https://www.getunleash.io/case-study-samsung) runs latency-sensitive bidding systems across many pods and data centers. Before flags, recovering from a bad change meant rolling a deployment back everywhere, with minutes of elevated errors while it propagated. Now the team disables a single flag and recovers in seconds, and its change failure rate has come down.

[The AA](https://www.getunleash.io/case-study-the-aa), often called the UK’s fourth emergency service, makes more than 95% of its sales digitally. Since adopting FeatureOps it ships 35% more deployments each month, year on year, and has had zero major incidents caused by a release.

## Dynamic configuration

A [variant payload](https://docs.getunleash.io/concepts/strategy-variants#variant-payload) in Unleash lets a flag carry JSON values your code reads at runtime. Add [constraints](https://docs.getunleash.io/guides/managing-constraints) and [segments](https://docs.getunleash.io/concepts/segments) on top, and you can change a timeout or a model name for one market without a CI/CD pipeline, with the audit log recording who changed it and when. We compared this approach with config files in [static config vs runtime control](https://www.getunleash.io/blog/static-config-vs-runtime-control).

### Configuration for AI agents

An agent runs on settings that change often: the model it calls, the prompt version, temperature, a token budget per conversation and the list of tools it may use. Keep them in a variant payload in Unleash and you can move an agent to a cheaper model when costs spike, or give a new prompt to 5% of conversations before everyone gets it.

When a new model came out, the Unleash team [moved its internal agent over](https://www.getunleash.io/blog/controlling-an-ai-agent-at-runtime-using-unleash) by editing one JSON value, which took about 15 seconds. The system prompt lives in the variants of a separate flag, so engineers and support staff each get a persona suited to their work. The agent settles on its model and persona once per Slack thread, and nobody sees either one change halfway through a conversation.

The tool list matters most when something goes wrong. If an agent starts acting oddly, you can cut it down to read-only tools in seconds and let it keep answering questions while the team works out what happened.

### Timeouts and retry budgets

Usually you configure timeouts and retry counts once, in code or a config file, and rarely look at them again. During a partial outage, those numbers decide whether your services back off politely or bury a struggling dependency under retries.

A payload per dependency with the timeout and the retry limit lets your SRE team tighten or loosen them while the incident unfolds. In Unleash, constraints on service name keep each change limited to the callers that need it.

### Quotas per API plan

A public API sold on several plans needs a quota for each one, plus the occasional exception, such as a higher limit for a customer during their own launch week.

Keep the quotas in a payload and define each plan once as a segment. The platform team raises the limit for one account with a constraint on its customer ID, and a [scheduled change request](https://docs.getunleash.io/concepts/change-requests#scheduled-change-requests) puts it back on the date the exception ends.

### Experiment winners that become settings

Say an A/B test of free shipping thresholds finds that $45 beats $35 and $60. The winning value can stay in the flag as the setting for everyone, where merchandising can lower it for the holiday season or set a different amount per market.

### Banners and incident messages

When a payment method goes down in one country, support wants a banner on that country’s checkout page within minutes. A payload holding the message text, targeted by market and app version, lets the support team publish and edit it themselves.

Pair it with a scheduled change and the banner disappears when the maintenance window closes.

### Search and recommendation weights

Retail search ranking blends relevance, margin, stock level and freshness. During a clearance sale, merchandisers want overstocked items higher in the results for a week.

With the weights in a payload, they can adjust the blend for one category, watch conversion through [impact metrics](https://docs.getunleash.io/concepts/impact-metrics) and restore the defaults when the sale ends.

## Keeping switches ready

A kill switch nobody has flipped in a year is a guess. Put each switch on a drill calendar and write down what happened when you flipped it.

Decide in advance who may flip which switch. [Change requests](https://docs.getunleash.io/concepts/change-requests) with four-eyes approval suit anything that relaxes a control, and a break-glass role lets on-call engineers act alone when minutes count. [Prudential](https://www.getunleash.io/prudential-case-study) treats every production flag change as an audited change, and Unleash syncs those changes and their approvals to ServiceNow in the background, so developers never file a ticket.

Make sure the switch still works when the network is part of the incident. Backend SDKs evaluate flags locally and cache the last known state. [Unleash Enterprise Edge](https://www.getunleash.io/unleash-enterprise-edge) gives frontend and mobile SDKs a nearby endpoint and can run in each region or failure domain you care about, which keeps the controls reachable during the outage they’re meant to contain.

Unleash is the FeatureOps platform that companies including Wayfair, Samsung, The AA and Mintlify use to control how their software behaves in production. If you’re working through any of the patterns above, [we’d like to help](https://www.getunleash.io/plans/enterprise).
