· Charlie Holland · Architecture  · 13 min read

"Can You Build RAGs?" — Revisiting a Question from 2024

Eighteen months ago, a senior exec at a global consulting firm asked me whether we could build RAGs. His customers were banging down the doors. The question is quietly awkward now — and most people still asking it don't actually need what they think they're asking for.

Eighteen months ago, a senior exec at a global consulting firm asked me whether we could build RAGs. His customers were banging down the doors. The question is quietly awkward now — and most people still asking it don't actually need what they think they're asking for.

Eighteen months ago I was in a room with a senior exec at one of the big global consulting firms. He leaned across the table and asked, with the tone of a man who had already decided the answer was yes:

Can you build RAGs? We’ve got customers banging down our doors for RAG and GenAI, and we just can’t build them fast enough.

It was late 2024. Vector databases were raising at valuations that made venture partners weep with joy. Pinecone had just closed a $100M Series B. Every enterprise architect in the Fortune 500 had a slide deck titled Our GenAI Roadmap, and bullet point three was always “RAG over our internal knowledge base.” LangChain had become a verb. You weren’t just building chatbots, you were chaining them.

Customers wanted it. Firms like his were hiring as fast as they could, at rates that implied a three-letter acronym was worth its weight in billable hours.

Eighteen months on, has it cooled?

The demand for “LLMs that know our data” is higher than it was. The specific architecture people used to mean by RAG has had most of the air let out of it, and a lot of what was sold as RAG was wishful thinking dressed up in cosine similarity.

Before we get into why, it’s worth remembering what RAG was solving in the first place.

Why we ever needed RAG

Take a frontier LLM from late 2023 — GPT-4, Claude 2, the first generation that felt properly useful. What could it actually do for your business?

It had been trained on the public internet, a shelf of books, and a mountain of GitHub code. It had no idea what your company sold, who your customers were, or where the leak in your 2024 Q3 revenue forecast was hiding. It had a training cutoff somewhere in the past, so anything that happened since was invisible to it.

Worse, it was stateless. Every conversation started from zero. Nothing persisted between sessions. Nothing from your Confluence, your Salesforce, your data warehouse. You could tell it about your business in the current prompt, but when you came back tomorrow it had forgotten everything you told it yesterday.

And the prompt wasn’t big. GPT-4 launched with an 8k context window. The 32k version was a paid upgrade. Claude 2 offered 100k and that felt gigantic. You couldn’t just paste a 300-page employee handbook in and ask questions — it wouldn’t fit, and if by some contortion it did, the model would lose track of things in the middle.

So you had three problems stacked on top of each other:

  1. The model knows nothing specific about your business.
  2. It has no memory between calls.
  3. The place where you could put business context — the prompt — was tiny.

RAG solved all three with one mechanism. Keep your business data outside the model, in a searchable store. When a user asks a question, retrieve the handful of passages most relevant to the question, and paste them into the prompt. The model “knows” about your business for the duration of that one call. No fine-tuning, no retraining, no waiting for OpenAI to index your Confluence.

It was an elegant workaround — and, given the constraints of 2023, close to the only one.

How naive RAG actually works

RAG stands for Retrieval-Augmented Generation. Strip it to the essentials and it’s three steps:

  1. Chop your documents into chunks.
  2. Embed each chunk into a vector using a model like OpenAI’s text-embedding-3-large or Cohere’s embed-v3.
  3. When a user asks a question, embed the question, find the k nearest chunks by cosine similarity, stuff them into the prompt, ask the LLM to answer using only those chunks.

The minimum viable RAG, in about fifteen lines:

from openai import OpenAI
import numpy as np

client = OpenAI()

def embed(text: str) -> np.ndarray:
    resp = client.embeddings.create(
        model="text-embedding-3-large", input=text
    )
    return np.array(resp.data[0].embedding)

def retrieve(query: str, chunks: list[dict], k: int = 5) -> list[dict]:
    q = embed(query)
    scored = [
        (np.dot(q, c["embedding"]), c) for c in chunks
    ]
    scored.sort(key=lambda x: -x[0])
    return [c for _, c in scored[:k]]

def answer(query: str, chunks: list[dict]) -> str:
    context = "\n\n".join(c["text"] for c in retrieve(query, chunks))
    prompt = (
        f"Use only this context to answer:\n\n{context}\n\n"
        f"Question: {query}"
    )
    resp = client.chat.completions.create(
        model="gpt-4", messages=[{"role": "user", "content": prompt}]
    )
    return resp.choices[0].message.content

That’s the thing that built a billion-dollar industry. LangChain existed to glue these steps together into “chains.” Vector databases existed to make step 2 fast at scale. Consulting firms existed to tell you which embedding model to use and how to chunk your PDFs.

In demos, it was magical. A user typed a question about HR policy, the LLM produced a coherent answer that cited a specific paragraph from the 2019 handbook. Executives clapped. Deals were signed.

Then it went to production.

Why naive RAG broke

The problems were easy to miss until real users asked real questions over real corpora.

Embeddings don’t know what you mean

Embedding models are trained on a lot of generic text. They capture lexical and topical similarity pretty well. What they don’t capture is the actual question you’re asking.

A question like “what’s our refund policy for enterprise customers?” will cheerfully retrieve chunks about consumer refunds, vendor refunds, expense refunds, and every other document that mentions the word “refund.” The chunk containing the actual enterprise refund policy — which uses the phrase “SLA credit” because the legal team rewrote it in 2022 — won’t show up. Retrieval quality depends entirely on how well the corpus vocabulary matches the question vocabulary, and when they diverge, you get confident nonsense.

The classic fix is hybrid retrieval — combine embeddings with BM25 keyword scoring and a reranker like Cohere Rerank or a cross-encoder. It works, but now your “RAG” is a three-stage pipeline with three models and three sets of failure modes.

I spent most of the 2010s at Smartlogic doing exactly this — minus the LLM at the end. Semaphore was a classification and enterprise search product: take unstructured documents, classify them against a taxonomy, index them with BM25, boost retrieval with semantic facets, rerank the top results, and hand them to the user. Some of the world’s largest pharma, media, and government organisations paid good money for that stack because BM25 on its own wasn’t enough and embeddings (as we now know them) didn’t exist yet. The classification layer was what made retrieval trustworthy at enterprise scale.

Watching the industry rediscover all of this from first principles in 2023 under the banner of “RAG” was like watching someone invent Elasticsearch because they hadn’t heard of Elasticsearch. A decade of enterprise search lessons — synonyms, precise terms, rerankers, taxonomy, metadata — had to be relearned by every vector-database startup with a demo. They mostly got there. It just took a year longer than it needed to.

Chunking is lossy and the boundaries matter

You split your documents. How? Fixed size? Sentence boundaries? Recursive? Semantic chunking?

Every strategy loses something. Consider a compliance document that opens with “Under no circumstances should the following claims be made to customers, as they expose the company to material legal risk:” and then cheerfully lists thirty things — “guaranteed 10x returns,” “zero risk of data loss,” “will not be outsourced to third parties,” “fully compliant with every jurisdiction we operate in.” Chunk that document at the wrong boundary and the retriever serves up the list with the warning neatly severed from it. The LLM, now holding a well-formatted catalogue of plausibly useful statements with no context explaining that they are in fact a list of things that will get you sued, proceeds to reassure the customer.

Getting the preamble reattached to the list at that point is going to take more than the silver tongue of one Jimmy McGill.

Engineers spend a remarkable amount of time tuning chunk sizes and overlap windows, and each tweak improves some queries and degrades others. It’s whack-a-mole, with k and chunk_size as the mallets.

The context is a sealed box

The LLM only sees the chunks you retrieved. If the retriever missed the relevant passage, the model has two choices: say “I don’t know,” or hallucinate. In practice, enterprise RAG systems hallucinate more than the marketing suggests, because the chunks look relevant — they mention the right keywords — so the model confidently synthesises an answer that’s subtly wrong. Plausible, cited, incorrect: the worst failure mode you can build.

Embeddings drift

You indexed the corpus in March with text-embedding-3-small. In November, OpenAI deprecates it. You re-embed everything with text-embedding-3-large. Retrieval quality changes overnight, better on some queries, worse on others. Your ground-truth eval sets were built on the old embeddings. Your prompts were tuned for the chunks the old retriever returned. Congratulations, you have a miniature migration project.

What changed

Between late 2024 and now, four things happened roughly at the same time.

Context windows grew to absurd sizes

Gemini 2.0 landed with 2M tokens of context. Claude hit 1M. GPT-5 runs at several hundred thousand. The original problem RAG was solving — “your document is too big for the prompt” — stopped being a problem for most documents.

A 300-page employee handbook is about 250k tokens. That fits in Claude’s context window with room to spare. You don’t need to chunk it, embed it, or retrieve from it. You send the whole thing.

The lost-in-the-middle problem still exists — models attend more to the start and end of their context than the middle — but it’s a much gentler failure than naive RAG’s “we lost the passage entirely.” And modern models are measurably better at long-context attention than they were in 2023.

Prompt caching changed the economics

This is the one that gets underestimated. Anthropic’s prompt caching and Google’s equivalent let you pay the full cost of tokens in the cached prefix once, and then pennies on subsequent reads.

import anthropic

client = anthropic.Anthropic()

# The whole handbook is cached on first use, then cheap thereafter
response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": EMPLOYEE_HANDBOOK,  # 250k tokens
            "cache_control": {"type": "ephemeral"},
        }
    ],
    messages=[{"role": "user", "content": "Enterprise refund policy?"}],
)

For a handbook that 500 employees query 50 times a day, the cached-prefix cost over a week is a rounding error. It’s often cheaper than running a vector database, paying for embeddings, and maintaining the retrieval pipeline that serves the same queries.

Around Q2 2025 this math flipped for the average enterprise corpus. “Just put the whole thing in the prompt” stopped being a naive strawman and became the default for small-to-medium knowledge bases.

Tool calling became reliable

In 2024, function calling worked sometimes. Models invented arguments, forgot to call tools, or called them in the wrong order. You couldn’t stake a production architecture on it.

By mid-2025 the frontier models had become genuinely good at multi-step tool use. Give the model a search_docs(query) tool, a run_sql(query) tool, and a read_file(path) tool, and it will plan a sequence of calls to answer a question. It decides when to search, what to search for, and whether to refine the query based on what it found.

That’s more powerful than one-shot retrieval. The model isn’t limited to whatever your embedding-based retriever thought was relevant. It can ask follow-up queries. It can cross-reference. It can say “that didn’t have what I needed” and try something else.

tools = [
    {
        "name": "search_docs",
        "description": "Search internal docs by keyword or phrase",
        "input_schema": {
            "type": "object",
            "properties": {"query": {"type": "string"}},
        },
    },
    {
        "name": "run_sql",
        "description": "Execute a read-only SQL query against the data warehouse",
        "input_schema": {
            "type": "object",
            "properties": {"query": {"type": "string"}},
        },
    },
]

# The model decides what to call, in what order, and when to stop

When your data is queryable (SQL, Elasticsearch, a filesystem, a REST API), tool calling beats embedding-based retrieval on accuracy, auditability, debuggability, and cost. Most of the time.

MCP turned tools into a protocol

Model Context Protocol standardised how agents discover and invoke tools. Boring infrastructure, which is why it matters. A tool your internal search team builds for one agent is available to any other agent, with no bespoke glue code in between. “Chain everything together” used to be a value proposition. Now it’s a protocol.

When is RAG still the right answer?

Not never. Specifically:

  • Genuinely huge corpora. If your knowledge base is tens of millions of documents and hundreds of billions of tokens, it doesn’t fit in a context window. Retrieval is the only way to narrow the problem.
  • Auditable and deterministic retrieval. In regulated domains (legal, clinical, financial) you often need to prove exactly which document the answer came from. A retrieval step returning a ranked list of passages is easier to audit than an agent that makes four tool calls and synthesises an answer.
  • Hard cost ceilings. At massive query volumes against a fixed corpus, retrieving five chunks and passing 5k tokens is still cheaper than prompt-caching a 500k-token corpus. The break-even point has moved, but it exists.
  • Latency-sensitive paths. One embedding lookup and one LLM call is faster than an agent that makes three tool calls sequentially. Sub-second responses favour the old pattern.

RAG has gone from “the default architecture for any LLM-plus-knowledge use case” to “one option in a toolkit, chosen when the constraints justify it.”

Would LangChain exist without RAG?

Almost certainly not, at least not in its original form. LangChain’s 2023 value proposition was chaining — compose retrievers, prompts, memory, and output parsers into pipelines. That was genuinely useful when the ecosystem lacked primitives and the model was the dumbest component in the stack.

The model is now the smartest component. Tool calling is native to the SDK. Caching is a flag in the API call. Most of what LangChain abstracted has either moved into the model provider SDKs or become a one-liner with native tool use. The chains dissolved into the models.

Do we just like RAG because it’s a TLA?

A little, yes. Three-letter acronyms confer seriousness. RAG is easier to put on a slide than “the LLM will call search_docs and run_sql as needed, cache the stable context, and synthesise an answer from what it finds.”

There’s a more interesting reason it stuck though. RAG gave organisations a mental model for how LLMs could use their data safely: the model only knows what we retrieved for it. For a security team, that’s a reassuring story. It’s also, as it turns out, mostly wrong — the model hallucinates around gaps in retrieval and cites chunks that don’t support its answer. But the story was useful even when the implementation was fragile.

The newer mental model is that the model is an agent searching your systems with scoped permissions. Harder to explain in a boardroom. Which is why “RAG” will stay in the vocabulary for another couple of years while the architecture underneath it continues to evolve.

Back to the exec

If the same exec asked me the question today — “Can you build RAGs?” — I’d say yes, but almost none of your customers need one in the way they’re asking for one.

What they usually need is:

  1. An LLM with tool access to their existing systems (SQL, search, APIs).
  2. A prompt-caching strategy for the stable parts of their context.
  3. Evaluation infrastructure to catch regressions when models or prompts change.
  4. A genuine retrieval pipeline — embeddings, hybrid search, reranking — for the slice of the corpus that actually demands it.

That’s a more boring answer. It takes longer to explain, and you can’t sell it as a TLA. But it’s the architecture that ships and holds up, and eighteen months after that meeting it’s the one customers are really asking for. Even if they’re still using the old word to describe it.

Back to Blog

Related Posts

View All Posts »
How Do You Contain a Thing That Knows How to Escape?

How Do You Contain a Thing That Knows How to Escape?

Agentic AI on your laptop is one thing — worst case it trashes your machine. But enterprises want it in the cloud, at scale, pointed at everything. The problem isn't the power. It's the containment. And this thing is smart enough to read the blueprints of its own cage.

Claude Is Not Your Architect. Stop Letting It Pretend.

Claude Is Not Your Architect. Stop Letting It Pretend.

Somewhere between 'ask Claude for a quick opinion' and 'Claude is writing our Jira tickets,' we lost the plot. AI agents are brilliant implementers. They're also confidently wrong about every decision that matters. And when it all falls over, they won't be the ones carrying the bag.

Claude Code Is Brilliant. It's Also Forrest Gump.

Claude Code Is Brilliant. It's Also Forrest Gump.

AI agents will get you to dev-done at record speed — right up until you realise you're miles past the actual problem. 80% of projects still spend 80% of their time at 80% complete. Agents don't fix that. They just help you reach it sooner.

AI Is a Model of Reality. It's Not Reality.

AI Is a Model of Reality. It's Not Reality.

Your exec promised a customer that AI would 'just know' the answer to any question across every database in the org. It won't. AI is a model of reality — and there's always a gap. The question is whether you know what it is.