Controlling an AI agent at runtime using Unleash
You probably already have a flag system and you are looking to experiment with and control your AI agent at runtime. Vendors are selling AI config or LLM management tools but you can already do a lot with existing tooling. At Unleash, we are developing our internal agents, like many of you. In this post I’ll show how we use Unleash to add an automated kill switch to safeguard our infrastructure spend, experiment with different models and effort levels, personalise the agent persona by targeting different users with strategy variants, and provide them with a consistent experience by picking the correct stickiness.
One agent, three flags
We interact with our agent in Slack. It’s built on the Claude Agent SDK, and our harness is written in TypeScript, so we started by adding the Unleash SDK for Node.js. As we began rolling out the agent to more users, we worked out what we wanted to act on, what we wanted to experiment with, and what we wanted to track. That came down to three feature flags:
- A kill switch that turns the agent off. With a safeguard added, we can automatically disable the agent when costs exceed a threshold.
- A model and effort flag, whose variant is a simple JSON payload.
- A persona flag, whose variant is the system prompt as plain text.
On top of those, we use impact metrics to track cost and satisfaction rating, and the Unleash SDK’s impression data to record which variant each person gets.
Killing the agent if things go wrong
The kill switch lets us stop both the agent taking any new queries and messages leaving its runtime. When the kill switch is disabled, everything is working as intended. When I enable the kill switch, the agent can only send error messages.
The Claude Agent SDK also exposes OTLP metrics, which let us gain insights into the cost per turn for each model. Then we hooked it up to our VictoriaMetrics instance, registered as an external impact metrics provider. This way, we can set up a safeguard in Unleash to disable the agent entirely if costs in the last three hours exceed roughly 10x the cost of a normal three hour period. We are using a high number that would clearly indicate something is wrong. This will help us sleep better for sure.
Trying out different models
Variants in Unleash let us serve different types of agent configuration at runtime, without a redeploy. We set up a flag for the model and effort with a single JSON variant:
With this setup, we easily updated to Opus 5.5 as soon as it came out. It took us 15 seconds to swap the value in the Unleash UI! Next, we are planning to run an A/B test with two variants, one with high effort and one with a lower effort, keeping an eye on the user satisfaction and the cost associated with each variant. This will let us optimise our agent costs and user experience going forward.
Different personas for different people
We also wanted to personalise the agent, so that different groups of people get a response style that suits them. The persona flag’s variant contains the system prompt as plain text, so a new persona is a new variant: no code change, no deploy.
Side note: our system prompts include multiple lines, and the string variant input in Unleash used to be a single-line field that dropped line breaks when you pasted one in. One perk of working on the tool you use: I made the field multi-line, so editing became a lot easier.
We currently run two personas: one for engineers and one for our support staff. The engineering variant goes into more detail and uses the tone an engineer expects, while the support variant simplifies topics a bit so we can provide the best support. In practice, we target users via the Slack user ID and segments in Unleash. When the flag is off, the agent falls back to its default persona, so turning it off is always safe. The satisfaction tracking below tells us whether each group is happy with its persona.
Consistent per user, fixed per thread
We believe it is important that everybody has a consistent experience with the agent, and since we use threads in Slack, two things gave us that. The first is stickiness on the Slack user ID. We provide the Slack user ID in the Unleash context, which is the default stickiness, so each person lands in the same bucket every time. As long as we don’t change the split, everybody gets the same variant in every thread.
The second is deciding once per thread. The model, effort and persona are evaluated when a new thread is created, and that variant stays in the thread forever. If we change a person’s assignment, they see the change once they start a new thread with the agent, never halfway through a conversation.
Tracking and reacting to satisfaction
We have our flags set up and the ability to change the model and persona at runtime. But what about a way to see how our changes or experiments are tracking over time?
This is where impact metrics come in. With the Unleash SDK for Node.js we can easily add Prometheus-style counters in the agent harness. We hook into the Slack reaction event and add two counters: one for thumbs up 👍 and one for thumbs down 👎.
When users react to agent messages, the reaction handler in the harness increments the impact metrics counter. This appears in the Unleash UI within seconds, giving us a quick glance at the overall satisfaction.
Next, we are planning to build out smarter evals and better ways to track satisfaction, as we already discovered that sometimes users find simple thumbs up and down too ambiguous.
What’s next
We are continuing to learn from this setup, while using Unleash as part of our agents’ infrastructure. Once the new evals produce a score, we’ll be able to feed it back into Unleash as an impact metric and put a safeguard on it, the same way we did with cost.
We also have two experiments lined up. The first is the effort A/B test I mentioned earlier. The second is splitting each group between its own persona and the default, to see whether tailoring actually helps. Both use the same user ID stickiness, so each person stays on one side of the split. Additionally, we’ll also tag our cost metrics with the persona in order to see the cost of each persona.
Start with what you have
You probably already use a feature flag system. With Unleash, we set our agent to switch itself off if costs spike, upgraded to Opus 5.5 in 15 seconds, assigned different personas, while keeping every conversation consistent without adding yet another tool.
If you only do one thing, put a kill switch on your agent before you need it.
