Skip to content
Owais Barkati

Federated fine-tuning of Llama 3 with Flower and QLoRA

Federated learning moves the model to the data instead of the data to the model. That inversion is only practical for large language models because QLoRA shrinks both what each client must hold in memory and what it must send back. Here is why the two techniques belong together.

Federated learning and QLoRA solve two halves of the same problem, and neither is sufficient alone. Federated learning lets a model train on data that cannot be centralised. QLoRA makes training a large language model cheap enough — in memory and in bandwidth — that doing it on distributed client hardware is actually feasible. Together they make private fine-tuning of a model like Llama 3 practical rather than theoretical.

I built this pairing into Sentinel, a multi-agent system where the training pipeline fine-tunes Llama 3 and Qwen across distributed client nodes using Flower for orchestration and QLoRA for the training itself. This post is about why that combination is the right shape, and where it gets uncomfortable.

The problem federated learning actually solves

Conventional fine-tuning has a precondition that often goes unexamined: the data has to be in one place. For public corpora this is uninteresting. For the data that would make a model genuinely useful — clinical records, transaction histories, internal documents, anything under GDPR, HIPAA or India’s DPDP Act — it is frequently the thing that cannot happen. Not “expensive to arrange.” Cannot.

Federated learning inverts the direction of travel. Rather than moving data to the model, it moves the model to the data. Each client node trains locally on records that never leave it, and only the resulting parameter updates are sent back and aggregated into a new global model. The server sees updates. It never sees records.

That inversion is the entire value proposition, and it is also where the cost appears.

Why plain federated learning falls over on an LLM

Classical federated learning was developed on small models — think next-word prediction on a phone keyboard. Apply the same algorithm to a modern LLM and two constraints bite immediately.

Memory. Full fine-tuning requires holding the model weights, the gradients, and the optimiser state. For Adam that is roughly four times the model’s parameter memory before activations. Client nodes in a federated setup are not H100 clusters; they are whatever the participating institution happens to have. Full fine-tuning of an 8B model is simply out of reach on most of them.

Bandwidth. Federated learning is iterative. Every round ships model updates from each client to the server and the aggregated result back. If an update is the full weight delta, each client is sending gigabytes, every round, and the network becomes the bottleneck long before compute does.

Both constraints come from the same source: treating all the parameters as the thing being trained.

What QLoRA changes

QLoRA attacks both at once by changing what “training” means.

The base model is quantised to 4-bit and frozen. It is never updated, so no gradients or optimiser state are needed for it, and its memory footprint drops roughly fourfold. Then low-rank adapters are injected into the model’s linear layers, and only those adapters train. Because the adapter matrices are low-rank, they hold a tiny fraction of the parameters of the layers they modify.

The federated consequences are direct, and they are the reason the combination works:

  • Memory fits on commodity client hardware, because the frozen quantised base needs no gradient or optimiser state.
  • The payload collapses. Clients exchange adapter weights, not full model deltas. Rounds that would have been gigabytes become megabytes, which is what turns the network from a blocker into an ordinary cost.
  • Aggregation gets simpler. Every client starts from an identical frozen base, so the only thing that differs — and the only thing worth averaging — is the adapters.

That last point is easy to skip past and is doing real work. Shared frozen weights mean the aggregation step operates on a small, uniform object rather than on the entire model.

Where Flower fits

Flower (flwr) handles the orchestration: client lifecycle, round coordination, and the aggregation strategy. Its useful property is that it is agnostic about your training code — a client exposes fit and evaluate, and what happens inside is ordinary PyTorch. A QLoRA training loop drops in without being contorted into a federated-specific framework.

The thing to get right is the boundary between Flower’s parameter exchange and the adapter weights. Flower moves arrays; the client is responsible for extracting only the adapter parameters and loading them back. If you hand it the full state dict, you have reintroduced exactly the bandwidth problem QLoRA just eliminated.

The part that stays hard: non-IID data

Here is what no amount of parameter-efficiency fixes.

Classical federated averaging assumes clients hold roughly similar data distributions. Real federated deployments violate this constantly — that is usually why the data is federated. Two hospitals see different patient populations. Two branches serve different demographics. The data is non-IID (not independent and identically distributed), and that is the normal case, not the edge case.

Under non-IID data, local training pulls each client’s adapters toward its own distribution. Averaging adapters trained in genuinely different directions can produce a global model worse than any individual client’s, and the symptom is a global metric that improves for a few rounds and then plateaus or degrades while local metrics look fine.

This is the honest limitation, and the thing to measure first. Evaluate the global model on each client’s held-out data separately rather than only in aggregate — an average score hides the case where the model got better for the largest participant and worse for everyone else.

Why this combination is worth the trouble

Federated learning alone is impractical for LLMs. QLoRA alone does not address privacy. Together they describe a setting that is otherwise unreachable: fine-tuning a capable open model on sensitive, distributed data, on hardware institutions already own, without any party surrendering its records.

That is a narrow problem. It is also the exact problem that stops a great deal of valuable data from ever improving a model.