Skip to content
Owais Barkati
AI / ML

The assistant on this site has no vector database

Why retrieval was the wrong answer for an 8,000-token corpus, and what replaced it

Role
Designed and built the assistant
Period
Cost per cached turn
~$0.0012

Claude Haiku 4.5, cache read at 0.1x input

Corpus size
8.8K tokens
Vector database
None

Deliberately

The decision worth defending

Almost every “chat with my site” build reaches for a vector database. Embed the content, store the vectors, embed the query at request time, retrieve the top k chunks, put those in the prompt. It is the default architecture, and for this site it would have been wrong.

The entire corpus — profile, five case studies, two papers, the posts — is about 8,800 tokens. A 200K-token context window holds it twenty-two times over. Retrieval exists to choose what to send when you cannot send everything. Here, you can.

Adding a vector store would have bought nothing and cost four things: a second system holding a copy of content that already lives in the repo, an embedding API call on every query, 100–300 ms of added latency, and a sync problem — the moment the index lags the content, the assistant starts answering from a version of the site that no longer exists.

What it does instead

The whole corpus goes into the system prompt behind an Anthropic prompt-cache breakpoint:

system: [
  { type: 'text', text: PERSONA },
  {
    type: 'text',
    text: `REFERENCE MATERIAL\n\n${KB_CORPUS}`,
    cache_control: { type: 'ephemeral' },
  },
]

Cache reads bill at 0.1× input. On Claude Haiku 4.5 at $1 per million input tokens, a cached turn is roughly $0.0012 — about twelve cents per hundred conversations.

The corpus is not written by hand. A prebuild step compiles it out of the same MDX and typed data files the pages render from, so the assistant and the site cannot disagree. The same bytes are also published at /llms-full.txt, which means the public machine-readable dump and the model’s knowledge are the same artefact rather than two things that drift.

Two constraints that shaped the build

Haiku 4.5 will not cache a prefix shorter than 4,096 tokens — the highest minimum of any current model. Below it, caching silently does not engage: no error, no warning, just a bill ten times larger than expected. So the build asserts a floor of 5,000 tokens and fails if the corpus drops under it, and a check on usage.cache_read_input_tokens confirms reads are actually happening rather than assuming they are.

The corpus must be byte-identical between builds. Caching is a prefix match on exact bytes. A build timestamp, a build ID, or unstable object key ordering changes the prefix, and every request pays full input price. The serializer therefore sorts collections by id, sorts keys, strips MDX machinery to plain text, and embeds no dates. Continuous integration hashes the generated corpus twice per run and fails the build if the two hashes differ — cheap insurance against a regression whose only symptom is a larger invoice.

Keeping a public LLM endpoint from becoming a bill

A chat endpoint on a public site is an open invitation to spend someone else’s money. The protections are layered, because any single one of them can be worked around:

  • Netlify’s native per-function rate limit, aggregated by IP and domain
  • An Origin allowlist, so the key only works for this site
  • A request-body cap and a server-side ceiling on conversation length
  • A bounded max_tokens, and Haiku rather than a frontier model
  • A spend limit on a dedicated API key — the only control that holds if every other one fails

The last of those matters most. Everything above it reduces the probability of a surprise; only a provider-side cap bounds the worst case.

What I would do differently

I would add an evaluation set before adding features. The assistant’s most important behaviour is refusing to answer what the corpus does not cover, and that is exactly the behaviour you cannot confirm by trying it a few times yourself. A fixed set of questions — some answerable, some deliberately not, some adversarial — scored on whether it declined when it should have, would turn “it seems to behave” into something measurable.

I would also revisit this architecture at around 100K tokens of content. The decision here is not “vector databases are overkill”; it is that this corpus is small. The point is to size the mechanism to the problem, and to be able to say why.

Stack

  • Claude API
  • Prompt caching
  • Astro
  • Netlify Functions
  • Web Streams