Sentinel: an autonomous multi-agent AI system
Intent-routed specialist agents, with federated fine-tuning so the data never leaves
- Role
- Designed and built the system
- Period
- Specialist agents
- 4
- Fine-tuning
- Decentralised
- Models
- Llama 3, Qwen
SQL, forecasting, RAG, web search
No raw client data leaves the node
Read this diagram as text
A natural-language query enters Sentinel and reaches an intent classifier before any answer is generated. LangGraph then dispatches the query along a declared edge to exactly one specialist agent. The SQL agent translates the question into a query against a live database, so analytical questions are answered from current data. The forecasting agent uses Facebook Prophet, a purpose-built statistical model that handles seasonality and trend explicitly, rather than asking a language model to extrapolate a series without error bars. The RAG agent answers from internal PDFs and policy files using FAISS vector stores with Hugging Face embeddings, grounding responses in source material. The web search agent covers anything outside the database, the document corpus and the model's training cutoff. Routing is the decisive stage because a misrouted query cannot recover downstream: 'how many loans defaulted last quarter' is a SQL question and 'how many will default next quarter' is a forecasting question, and they differ by one word. Making the router an explicit, inspectable graph node rather than emergent behaviour of a larger prompt means a misroute is a traceable edge instead of an opaque judgement.
What Sentinel is
Sentinel is an autonomous multi-agent framework that treats “answer this question” as a routing problem before it is a generation problem. A natural-language query arrives, an intent classifier decides what kind of question it actually is, and LangGraph dispatches it to the agent equipped to answer: a SQL agent, a forecasting agent, a retrieval-augmented generation agent, or a web search agent. FastAPI fronts the system.
The premise is that a single general-purpose agent handed every tool is the wrong shape. Tool choice degrades as the toolset grows, failures become hard to attribute, and the prompt carries instructions for situations the current query will never touch. Specialist agents behind a router keep each agent’s context narrow and its failure mode legible — when a forecasting answer is wrong, the forecasting agent is wrong.
Why routing is the hard part
Routing is where this design earns or loses its advantage, because a misrouted query cannot recover downstream. “How many loans defaulted last quarter?” is a SQL question. “How many will default next quarter?” is a forecasting question. They differ by one word, and the agent best suited to each shares no machinery with the other.
Making the router a distinct, inspectable stage rather than an emergent behaviour of a larger prompt means routing decisions can be examined and corrected on their own terms. LangGraph’s explicit graph structure is what makes that possible: the control flow is a declared graph rather than whatever the model decided to do this time, so a misroute is a traceable edge rather than an opaque judgement.
The four agents
The SQL agent translates natural language into queries against a live database, so analytical questions are answered from current data rather than from whatever ended up in a model’s training set.
The forecasting agent uses Facebook Prophet for time-series prediction. This is a deliberate rejection of the reflex to make the language model do everything: Prophet is a purpose-built statistical model that handles seasonality and trend explicitly, and an LLM asked to extrapolate a series produces plausible numbers with no error bars and no method.
The RAG agent answers from documents — internal PDFs and policy files — using FAISS vector stores with Hugging Face embeddings, so responses are grounded in source material rather than recalled.
The web search agent covers what the other three cannot: anything outside the database, the document corpus, and the model’s training cutoff.
Federated fine-tuning, and why it exists
The part of Sentinel I find most interesting is the privacy-preserving training pipeline, built with Flower and QLoRA to fine-tune Llama 3 and Qwen across distributed client nodes.
Conventional fine-tuning requires pooling data centrally — which is exactly what the institutions holding the most valuable data cannot do. Federated learning inverts this: the model travels to the data, trains locally on each node, and only the resulting parameter updates are aggregated. Raw records never leave the node that owns them.
That inversion has a cost, and QLoRA is what pays it. Full fine-tuning of a model the size of Llama 3 is out of reach on commodity client hardware, and shipping full weight updates between nodes each round would make the network the bottleneck. QLoRA quantises the base model and trains only low-rank adapters, which shrinks both the memory needed on the node and the payload sent for aggregation — the two constraints that otherwise make federated LLM fine-tuning impractical rather than merely difficult.
What I would do differently
I would build routing evaluation before adding the fourth agent. Every agent added expands the router’s decision space, and without a labelled set of queries and expected destinations there is no way to tell whether a new agent improved coverage or just made misrouting more likely. I would also make each agent return its confidence and its source alongside its answer, so the orchestrator can decline rather than confidently forward a weak result — in a system built around delegation, knowing when not to answer is the feature that makes the rest trustworthy.
Stack
- LangGraph
- FastAPI
- Flower
- QLoRA
- FAISS
- Hugging Face
- Prophet