Skip to content
Owais Barkati

LangGraph supervisor routing: getting a query to the right agent

One agent holding every tool degrades as the toolset grows. A supervisor that classifies intent first and dispatches to a specialist keeps each agent's context narrow and each failure attributable — but routing becomes the single point where the system can be wrong beyond recovery.

The reflex when building an LLM application is to give one agent every tool and let the model decide. It works at three tools. At eight it starts choosing badly, and when it chooses badly you cannot tell whether the tool was wrong, the arguments were wrong, or the question was never answerable by that tool at all.

A supervisor architecture inverts this. Classify what kind of question arrived, dispatch it to the agent built for that kind, and let that agent work with a narrow toolset and a prompt written for one job. I used this shape in Sentinel, routing to four specialists — SQL, forecasting, retrieval-augmented generation, and web search — with LangGraph holding the control flow and FastAPI in front.

Why the graph is the point

LangGraph’s contribution is not that it calls models. It is that control flow becomes a declared graph rather than whatever the model decided this time.

That distinction matters for debugging more than for architecture. In a single-agent loop, a bad answer is an opaque judgement somewhere inside a long prompt. In a supervisor graph, a bad answer has a traversal: this node classified the query as forecasting, that edge dispatched it, the forecasting node produced this. When the answer is wrong, you can say which edge was wrong. That is the difference between a system you can improve and one you can only re-prompt and hope.

The routing problem is the whole problem

A misrouted query cannot recover downstream. The specialist that receives it will answer confidently with the wrong machinery, and nothing later in the graph is positioned to notice.

The canonical example, from a loan-analytics context:

  • “How many loans defaulted last quarter?” — a SQL question. The answer is a fact sitting in a database.
  • “How many will default next quarter?” — a forecasting question. The answer does not exist anywhere; it has to be produced by a model with error bars.

They differ by one word. They share no machinery. And a system that answers the second with SQL returns an empty result set, while one that answers the first with a forecaster returns a prediction of the past.

This is why the router deserves to be an explicit, inspectable node rather than an emergent behaviour of a larger prompt. It is the stage where being wrong is unrecoverable, so it is the stage that should be easiest to examine.

What each specialist buys

The SQL agent turns a question into a query against live data, so answers reflect the database now rather than whatever was in the model’s training set.

The forecasting agent uses Prophet — a purpose-built statistical model that handles trend and seasonality explicitly. This is a deliberate refusal to make the language model do everything. Ask an LLM to extrapolate a time series and it will produce plausible numbers with no confidence interval and no method you can interrogate. For anything a decision rests on, “plausible” is the failure mode, not the goal.

The RAG agent answers from documents with a vector store and embeddings, so responses are grounded in source material rather than recalled.

The web search agent covers what the other three structurally cannot: anything outside the database, outside the document corpus, and after the training cutoff.

Each has a narrow prompt and a small toolset. None carries instructions for situations it will never see.

Where this gets harder than the diagram suggests

Routing accuracy is unmeasured by default. Every agent you add expands the router’s decision space, and without a labelled set of queries and expected destinations there is no way to tell whether a new specialist improved coverage or just made misrouting more likely. Build that set before the third agent, not after the fifth.

Some queries are genuinely multi-agent. “Show me last quarter’s defaults and project next quarter” is both branches. A router forced to pick one will pick one. Decide early whether your graph supports fan-out and a join, or whether you decompose first — adding that later means rewriting the supervisor.

Confidence has to travel with the answer. If each specialist returns only a result, the orchestrator has no basis to decline. If it returns a result, a confidence, and its sources, the orchestrator can say “I’m not sure” instead of confidently forwarding a weak answer. In a system built entirely around delegation, knowing when not to answer is what makes the rest of it trustworthy.

Streaming it through FastAPI

The architectural cost of a supervisor is latency: classification happens before any work starts, so the user waits through a decision that produces no visible output.

Streaming is what hides it. Emitting graph events over server-sent events as the traversal progresses means the interface can show that routing happened and which specialist took the query, while that specialist is still working. The wait becomes legible rather than blank — and legibility is most of what people mean when they say an interface feels fast.