Runtime control doesn’t have to be instant: you control “when”
Here’s a paraphrased objection that I hear frequently from infrastructure engineers when we discuss runtime control:
If I put my database connection pool behind a feature flag, will it yank the connection out from under a running query?
No. If you design it properly – and that reveals something important about how Layer 3 can work. Something the feature flag literature rarely addresses, because it focuses on the visible (like toggles affecting UI) and ignores the invisible cases (like back-end or infrastructure-level configuration).
Let’s start with an easy case first: a feature flag controls whether users see a new checkout flow. The evaluation happens per request, so the effect seems immediate. An ongoing request sees the old flow. The next sees the new flow. No state carries over, no cleanup needed.
Now a harder one: a feature flag controls part of the system that carries an underlying state. Think of an asynchronous process or a long-running transaction, where you can’t swap underlying code between requests without hard-to-predict negative consequences. The decision to switch is available instantly. The application of that decision must respect the resource’s lifecycle.
This is the problem of when. It’s also an underlying reason for the initial objection and the proper topic we need to address. Because where the decision lives and when its effect applies are two separate dimensions.
Four application cadences
Not every change can and should take effect at the same pace. Here are the four cadences, mapped to where they naturally fit.
Per-session (sticky, transactional)
The decision is made once per session (or per entity, per long-running transaction) and cached for the duration. Changing the flag mid-session would create inconsistency. For example, a user sees variant A on page load and variant B on the next click, or when a business operation requires an orchestrated chain of microservice calls.
# Per-session: evaluate once, cache for the session.
def get_session_features(user, session):
context = {"userId": user.id}
if not session.has("pricing_tier_variant"):
variant = client.get_variant("pricing-experiment", context)
session.set("pricing_tier_variant", variant)
return session.get("pricing_tier_variant")
Beyond pricing-tier experiments, per-session cadence also applies to tenancy gating or permission sets, personalization profiles, or authentication-flow variants.
Per-lifecycle-event (deferred)
The system knows the new desired state instantly. Still, it applies it at the next safe lifecycle boundary, e.g., during a connection pool drain, a cache TTL expiry, or a scheduled maintenance window.
# Per-lifecycle-event: flag is the intent register,
# pool drain is the application moment.
class ConnectionPoolManager:
def __init__(self, flag_client):
self.flag_client = flag_client
self.current_endpoint = "postgres://primary:5432/app"
def on_pool_drain(self):
"""Called when the pool naturally recycles connections."""
desired = self.flag_client.get_variant("db-endpoint", {})
if desired.value != self.current_endpoint:
self.current_endpoint = desired.value
# Safe: no in-flight queries.
self.rebuild_pool(self.current_endpoint)
Other examples are database migrations, connection pool rotation, certificate renewal, and cache strategy switch.
Per-request (immediate)
Each request evaluates independently. Changing the flag between the Nth request and the (N+1)th request has no side effects. This is also the FeatureOps sweet spot: full context, instant effect, zero deployment.
# Per-request: every evaluation is independent.
def handle_request(user, request):
context = {"userId": user.id}
if client.is_enabled("new-search-algorithm", context):
return search_v2(request.query)
return search_v1(request.query)
In addition to the above, this also applies to A/B experiments, kill switches, OpsToggles handling observability or log verbosity, chaos fault injection, or targeted profiling.
Per-restart (process-lifetime). The value takes effect only when the process starts. Think of application runtime environment configuration, like JVM heap size, the port the application listens on, loaded application modules, or dependencies. These are Layer 1 or Layer 2 decisions – changing them requires a service restart (if that’s a configuration) or even a new version of a built artifact (if the value is hardcoded).
The intent register pattern
The 3rd category (per-lifecycle-event cadence) is interesting for infrastructure engineers. The feature flag system becomes an intent register. It records that the desired state has changed instantly, and the application drains toward that state at its own pace.
In the most basic form, feature flag evaluation is an if statement in the code. And because it is code, it can be molded however you would like – including the when. Even if the flag evaluates to the new value immediately, it is an awareness at the control-plane level (inside a tool like Unleash that serves flag state and evaluation rules). As the system designer, you decide how to write the consuming code that decides when to apply it.
# The flag system decides WHAT (instant).
desired_state = client.get_variant("cache-strategy", {})
# The application decides WHEN (lifecycle-aware).
if desired_state.value != current_strategy and cache.is_drained():
# THIS is our safe moment.
apply_new_strategy(desired_state.value)
The code delegates what to the external system – that’s the Inversion of Control over activation authority from the previous posts. But it retains control over when to apply safely. The flag system provides the signal. The application provides the timing.
Conceptually, you can think of it as refueling an aircraft in the air. In software engineering, we have a similar idea called code reloading (yes, it is a thing: Erlang and its language runtime use it). With it, you can load a new code version, and it’s available for execution immediately. But existing processes don’t jump to the new version mid-execution. They keep running the old code until the runtime environment finds a safe point to switch to the new version at a natural execution boundary. The environment also manages two versions simultaneously. As a result, the decision to upgrade is instant. However, the application respects the process lifecycle.
Two dimensions of consistency
A second dimension matters enormously in distributed systems, and most feature flag discussions ignore it entirely: where in the stack does the evaluation apply consistently?
Temporal stickiness: user gets the same evaluation across time.
If user X is in cohort B for a pricing experiment, they must stay in cohort B across every request in their session – and potentially across sessions, if the experiment runs for weeks. Without temporal stickiness, a user sees $9.99 on the product page and $14.99 at checkout. Unleash handles this through consistent hashing on a stickiness key (userId or sessionId).
Cross-service stickiness: user gets the same evaluation across every service in the request chain.
This is the harder problem. In a distributed system, a single user request traverses an API gateway, a checkout service, a payment service, an inventory service, and an analytics pipeline. Or in business terms: from an email campaign to a marketing website and checkout system, ending at the shipping tracker. Throughout the whole customer journey, we must ensure pricing consistency. If user X is in cohort B at the API gateway, they must be in cohort B in every downstream service.
Without cross-service stickiness, a single request gets fractured across variants as it traverses the stack. The checkout service shows the new flow, but the payment service processes with the old logic. The experiment data is garbage because the cohort B population isn’t actually experiencing a consistent variant.
Request for user X enters the system:
WITHOUT cross-service stickiness:
API Gateway → evaluates flag → cohort B ✓
Checkout → evaluates flag → cohort B ✓ (lucky, same hash)
Payment → evaluates flag → cohort A ✗ (different seed)
Analytics → no flag SDK → no cohort ✗
WITH cross-service stickiness:
API Gateway → evaluates flag → cohort B ✓ → propagates context
Checkout → reads context → cohort B ✓
Payment → reads context → cohort B ✓
Analytics → reads context → cohort B ✓
This is why frontend-only A/B testing is easier, as the frontend is the only place where context and evaluation naturally coexist. The browser has the user identity, the session, and the SDK – all in one runtime. No context propagation needed.
You can solve that in Layer 3 with context propagation for both dimensions. Either through request headers, OpenTelemetry baggage, or SDK-level evaluation with shared context. But only if the feature flagging platform models them natively. A homegrown system that returns a boolean from a database has neither temporal stickiness (no consistent hashing) nor cross-service stickiness (no context propagation). It’s just a toggle, not a runtime control system.
The 2×2 diagnostic
Splitting by where and when across those two dimensions gives infrastructure engineers a practical tool. To externalize a configuration decision to Layer 3, ask those two questions:
Where does the decision live?
The answer determines the reversal speed:
- Layer 1: hours,
- Layer 2: minutes,
- Layer 3: seconds.
When should the effect apply?
The answer determines the application cadence:
- per-request,
- per-session,
- per-lifecycle-event,
- per-restart.

The Two-Question Diagnostic for Application Cadences
Examples: per-request decisions at Layer 3 are the sweet spot – all FeatureOps pillars operate at full power here. Per-lifecycle-event decisions at Layer 3 are the most sophisticated – the flag system is an intent register, and the application controls timing. Per-restart decisions belong at Layer 2 – no point externalizing something that can only change with a pod restart.
What homegrown systems miss
If you need a clear difference between a homegrown feature flagging service and a FeatureOps control plane, here it is: layering application cadence with gradual rollouts, targeting, and variants is rarely an option in a homegrown system.
They usually return an evaluation result, and the system designer must use it immediately or implement additional mechanisms to support advanced targeting with a more nuanced decision-application moment. As a result, responsibility shifts onto application teams that build increasingly complex application code around their simple flag system. All caching layers, session stores, lifecycle hooks, all custom and competing with feature delivery for engineering time. A mature platform models these semantics natively: stickiness keys, consistent hashing, advanced targeting, standardization around context propagation (like OpenTelemetry Baggage), variant evaluation with fallback defaults.
That’s the difference between a feature flagging service and a runtime control system – and it’s the difference that infrastructure engineers care about most, because they’re the ones who deal with the consequences when consistency breaks.
This is Post 5 of “The Runtime Control Layer” – a series on FeatureOps for infrastructure engineers.
Previously: GitOps Deploys Your Code. What Deploys Your Decisions?
Next: Blurred Lines: When Infrastructure Decisions Leak Into Runtime (and vice versa)
