Most LLM agents follow the same pattern: read the user’s query, call some tools, gather data, and respond with a data-aware answer. It works well in general, but support workflows need something this pattern doesn’t give you: every single issue has to follow the same resolution procedure each time it comes up.
Take order tracking, for example, a common request from the content creators who use Wishlink to share shoppable links and earn commission. When one asks whether an order was tracked via their affiliate link, the agent has to gather the order ID, brand, and purchase date, then search the database, then debug if needed. This must happen in that same exact order on every such ticket. If the agent decides on its own to skip a step or run them out of order, the ticket gets closed with a wrong answer.
Keeping this in mind, we built Wishii, the AI agent that handles creator support for Wishlink. Wishii runs across Freshchat (creator-facing support channel), Slack (for our Customer Success Managers), and our in-app chat. It covers the whole creator lifecycle, from onboarding to payouts, dozens of different ticket types, each with its own resolution procedure.
How Wishii is Wired
Before any design decisions make sense, here’s how the system looks. A single creator message runs through a fixed pipeline of steps: gather context, interpret the message, do the necessary lookups, and deliver a response. To build this pipeline we chose LangGraph, an open-source library for expressing LLM workflows as a directed graph. Nodes are functions that read and update a shared state object; edges decide which node runs next. If you’ve used a state machine, the model is familiar: nodes are states, and a node can branch to several others.
The core loop looks like this:
A turn runs through five nodes:
- LOAD_CONTEXT opens the turn by pulling the creator’s profile and chat history.
- ORCHESTRATOR is the one node that calls the LLM, to interpret the creator’s message.
- FLOW_POLICY is a deterministic, non-LLM node that decides what the conversation does next: deliver a reply, run more lookups, or loop back for another round of interpretation. We call this logic the engine; it’s the piece we keep referring to below.
- WORKER is a scoped subagent: an LLM that runs a bounded loop of tool calls to do the actual data lookups, then hands a structured result back to
FLOW_POLICY. (We cover how workers are bounded and scoped later.) - DELIVER closes the turn by sending the reply out to the channel.
Two edges are worth naming now, because they come back later. The loop-back edge (not done) means “interpret again”: the engine has decided this turn isn’t over. The done edge is any terminal reply: sending an answer, closing the ticket, and so on. Hold onto those; we’ll name them precisely in a moment.
What a “Flow” Actually is
Before we get to the interesting part, let’s set up some vocabulary we’ll use from here on.
- A flow is a scripted, multi-step procedure for handling one kind of request (think of a runbook a human agent would follow). Flows are an independent set, each with its own trigger, and the orchestrator starts whichever one matches the creator’s request.
- A slot is a single named piece of information a flow needs before it can act (think of a required field in a form). A slot gets filled one of two ways: the LLM reads it out of the creator’s message, or a deterministic regex extractor picks it up first and takes precedence. A flow collects its slots, does some work, and potentially starts another flow.
- A transition is a rule that hands control from one flow to the next once a condition is met.
Here’s a real flow that diagnoses “Why didn’t I get commission for this order?”, one of our most common queries. It’s defined as plain data:
ATTRIBUTION_DEBUG_COLLECT_FLOW = FlowDef(
id=FlowId.ATTRIBUTION_DEBUG_COLLECT,
slots=(
Slot(name="brand_name", description="The brand the creator purchased from ..."),
Slot(name="purchase_date", description="Approx date the order was placed ...", required=False),
),
transitions=(
Transition(trigger=SlotsFilled(), target=FlowId.ATTRIBUTION_DEBUG_FRESHNESS_CHECK),
),
)
Reading it from top to bottom: this flow needs a brand_name (required) and a purchase_date (optional). The single transition says: once the slots are filled (SlotsFilled()), move to the next flow, FRESHNESS_CHECK. That next flow checks whether the order is simply too recent to have arrived in our system yet; if it’s old enough to investigate, it moves on to collect an order_id and run the actual diagnosis. So one creator question walks through a chain of small flows, each collecting what it needs and passing control forward:
Two more pieces of vocabulary and we’re done with the jargon:
- A transition guard is a boolean condition on slot values that decides whether a transition is allowed to fire or not. The
COLLECTflow above had a single unguarded transition, but a later step in the same chain (the one that collects theorder_id) has two exits guarded on aquery_typeslot:
transitions=(
Transition(
trigger=SlotsFilled(),
target=FlowId.ATTRIBUTION_DEBUG_LOW_COMMISSION,
guard=SlotEquals("query_type", "low_commission"),
),
Transition(
trigger=SlotsFilled(),
target=FlowId.ATTRIBUTION_DEBUG_DIAGNOSE,
guard=Not(SlotEquals("query_type", "low_commission")),
),
)
Both transitions fire on the same trigger (SlotsFilled()); the guard is the tie-breaker. If query_type is "low_commission" we route to the low-commission flow; every other case, including zero or missing commission, falls through the Not(...) guard into the diagnosis flow, ATTRIBUTION_DEBUG_DIAGNOSE. Because guards compose (And, Or, Not), this routing is just data, not branching code you write by hand.
- Flows can pause and nest. Parked flows live in a stack: each parked flow is a frame that holds the flow’s collected slots. A frame is complete when all its required slots are filled.
That’s the whole model: flows made of slots, connected by guarded transitions, stacked when they nest. Keep the attribution example in mind. Every mechanism below acts on exactly this structure.
What Happens When a Creator Switches Topic Mid-Diagnosis?
Say a creator is halfway through an attribution diagnosis and suddenly asks why their Instagram account won’t connect. The model spots the topic switch and starts the social-login flow; the attribution flow parks where it was, slots intact, and resumes automatically once the new flow finishes. If the creator drops the original topic, the parked flow is discarded instead.
The Split: LLM Proposes, Engine Decides
The core architectural decision in Wishii is a hard separation between two things: understanding what the user means (language) and deciding what to do next (control flow). We put the LLM in charge of the first and explicitly refuse to let it own the second.
Concretely: the ORCHESTRATOR node calls the model once per turn and gets back a list of typed commands, not an action it executes directly.
Command = Annotated[Union[StartFlowCmd, SetSlotCmd, CorrectSlotCmd, CancelFlowCmd,
RespondCmd, DelegateCmd, EscalateCmd, ResolveCmd, ClarifyCmd], Field(discriminator="type")]
class OrchestratorSchema(BaseModel):
detected_language: str | None
reasoning: str
scratchpad_thought: str | None
commands: list[Command]
Note that OrchestratorSchema is a structured-output schema: instead of parsing the free text, we make the model return JSON that conforms to this exact shape. Most major LLM APIs support this, so commands arrive as validated objects.
Each command is a proposal, nothing more. The model can say “start the attribution flow” (StartFlowCmd), “set brand_name to Nykaa” (SetSlotCmd), or “this ticket is resolved” (ResolveCmd), but it never carries any of them out. It hands the list to the engine, and the engine decides what actually happens.
Here is how that looks in a single turn. Suppose the creator says “my Nykaa order didn’t get any commission.” The model emits a short, ordered list of commands:
[
StartFlowCmd(flow_id=ATTRIBUTION_DEBUG_COLLECT),
SetSlotCmd(flow_id=ATTRIBUTION_DEBUG_COLLECT, name="brand_name", value="Nykaa"),
]
The engine applies them in the order they arrive. Each command names its own target flow: the first starts the attribution flow, the second fills its brand_name slot. Once the list is applied, the engine reads the resulting flow state and produces exactly one outcome for the turn.
The engine can also overrule a proposal outright. If the model says the ticket is resolved but an escalation was raised in the same turn, the escalation wins and the ticket stays open. A flow can also name another flow as its prerequisite. If the model tries to start it before that prerequisite has finished, the engine drops the command. The model owns language; the engine owns control flow.
So what does the engine hand back? Exactly one Decision per turn, from a fixed set: RESPOND (send this text to the creator), CLARIFY (ask them to pick between a few options), DELEGATE (run a worker lookup), ESCALATE (hand off to a human), RESOLVE (close the ticket), or REORCHESTRATE.
That last one is the tell. REORCHESTRATE is the one Decision the model can’t propose; it isn’t in the command union at all. It’s the not done edge from the pipeline diagram earlier, a signal the engine sends back to the ORCHESTRATOR node meaning “loop again, you’re not done.” It fires in three cases:
- a new flow just started but its first required slot is still empty, so the model has to run again to ask for it;
- the model tried to delegate a lookup while a flow still has empty required slots;
- a
RespondCmdcame back with no actual text.
The LLM proposes; the engine is the only thing allowed to declare a turn over.
Workers: A Scoped Tool Loop, Not a Free-Roaming Agent
When the engine decides to DELEGATE, the work goes out to one or more WORKER nodes. We use LangGraph’s Send() API for this, which lets us spawn however many workers a turn needs at runtime, rather than wiring them into the graph ahead of time. So a turn that needs three lookups spawns three workers, and they run in parallel.
Each worker runs a bounded ReAct loop, the standard agent pattern where the model goes back and forth between reasoning and calling a tool, looking at each result, until it has enough to answer. We cap both the number of rounds and the calls per tool, so a confused worker can’t spin forever. Each worker also only sees the tools tagged for its own subtask, so a worker handling a payout question literally cannot reach an Instagram-login tool.
A worker returns a structured result: its findings, a summary, and a list of signals that feed back into the engine’s transitions. Recall the attribution flow: its freshness-check worker emits outside_freshness_window, which is exactly the signal the next transition looks for. This is the handoff back from the worker’s guesswork to the engine’s fixed rules.
Each flow lists its required_findings. We build a result schema per flow that rejects any worker output missing one of those keys, so a worker can’t hand back half an answer.
What Production Broke
The control-flow split held up fine. What broke was the stuff around it: the connection to OpenAI, and the bill it ran up.
The circuit breaker. Every LLM call went to OpenAI. One day the account hit its quota and Wishii stopped answering: a billing problem had become an outage. So we put a circuit breaker in front of the provider. When calls to the primary start failing on quota, the breaker opens and traffic switches to a fallback provider automatically. It fires one Slack alert when it opens, then closes again once the primary recovers. A quota failure on one provider no longer takes Wishii down.
Prompt caching. LLM providers bill a repeated prompt prefix at a fraction of the price, but only if that prefix is byte-for-byte identical. Ours wasn’t. The system prompt mixed fixed instructions with per-turn state like the date and the active slots, so every request looked new and nothing cached. We split the prompt into a static part and a per-turn part. The static part is now the same across conversations, so it caches. Today about 74% of the input tokens we send are served straight from that cache.
Where Wishii Landed
Wishii runs in production today across Freshchat, Slack, and our in-app chat. In a typical month it handles around 12,500 creator conversations, and about 95% of them go all the way through without a human stepping in. That last number is only comfortable to run because the engine, not the model, is what decides when a ticket can close. It costs us roughly 2-3 cents a conversation today, down from about 48 cents before we optimized the pipeline.
What We Learned
If you build an agent past the demo stage, a few things we’d pass on:
- Decide early which choices the model gets to make, and which it doesn’t.
- Let the model read language and suggest what to do. That’s the part it’s good at.
- Keep the decisions that determine a ticket’s outcome in plain code you can test on its own, without an LLM in the loop.