Many AI projects fail because RAG is chosen by reflex. Here are the questions to ask before deciding between RAG, classic search, or a plain API call.
Retrieval-augmented generation has become the default answer to almost every AI question in business meetings. Someone wants an assistant that knows the company's documents, and within minutes the conversation lands on vector databases, chunking strategies and embeddings. Sometimes that is the right answer. Often it is not.
RAG is a technique, not a goal. It means retrieving relevant passages from a knowledge base at query time and passing them to a language model so it can answer using that context. It solves one specific problem well: making a model answer from information it was never trained on, without retraining it. Before committing to it, it is worth checking whether that is actually your problem.
Start with the question, not the technology
Before choosing a stack, write down what users will actually ask. Not the topic — the questions. "What is our refund policy for enterprise contracts?" "Which supplier delivered late last quarter?" "How do I configure SSO in the internal tool?"
Then ask three things about each question:
- Does the answer exist in written form somewhere?
- Is it stable, or does it change weekly?
- Does finding it require reasoning, or just locating?
If the answers live in a handful of documents that rarely change, you may not need retrieval at all — you may need a well-structured prompt. If they live in a database with exact fields, you need a query, not a model. RAG earns its place when the knowledge is large, text-heavy, changing, and spread across sources.
RAG versus fine-tuning versus a plain API call
These three get confused constantly, and the confusion is expensive.
A plain API call with a well-written system prompt handles more cases than people expect. If your assistant needs a tone, a format, a set of rules and a small amount of fixed context, this is enough. It is cheap, fast and easy to debug.
Fine-tuning changes how a model behaves — its style, its format discipline, its domain vocabulary. It is a poor tool for injecting facts, because facts change and retraining does not. If your problem is "the model does not sound like us" or "it keeps ignoring our output format", fine-tuning is worth considering. If your problem is "it does not know our 2024 pricing", fine-tuning is the wrong lever.
RAG is for knowledge that is too large or too volatile to fit in a prompt and too factual to be learned by training. It keeps the source of truth outside the model, which is its main advantage: update the document, and the next answer reflects it.
A practical rule: use prompting for behaviour, fine-tuning for form, RAG for facts.
Where RAG genuinely pays off
RAG tends to work well in a few recognizable situations.
- Internal document assistants. Policies, procedures, technical documentation, meeting notes. The corpus is large, the questions are varied, and users accept an answer with a source link.
- Support and pre-sales. Answering from product documentation, contracts and past tickets, where the value is speed and consistency rather than creativity.
- Regulated or auditable contexts. When you must show *why* an answer was given, retrieval gives you a citation trail that a pure model call cannot.
- Fast-moving knowledge. Pricing, inventory rules, compliance updates. Anything where a stale answer is worse than no answer.
Notice what these have in common: the knowledge exists in text, users ask open questions, and traceability matters.
Where RAG quietly fails
RAG fails in predictable ways, and most of them are visible before you write any code.
It fails when the documents are bad. If your knowledge base contains five contradictory versions of the same policy, retrieval will faithfully surface the contradiction. Cleaning the corpus is unglamorous and usually more valuable than tuning the retriever.
It fails when questions need aggregation. "How many customers churned in Q3?" is not a retrieval question, it is a database question. A model reading ten document fragments will approximate an answer. A SQL query will be correct.
It fails when users expect precision on numbers, dates or calculations. Language models are not calculators, and retrieved text does not change that.
It fails when nobody owns the content. A RAG assistant is only as current as the documents behind it. Without a maintenance routine, quality decays within months.
A short decision checklist
Before starting, answer these honestly:
- Can a keyword search or a filtered database query solve this? If yes, do that first — it is cheaper and more reliable.
- Is the corpus clean, deduplicated and owned by someone?
- Do users ask open questions, or do they need exact figures?
- Do you need citations for compliance or trust?
- Can you measure quality? Define a small set of test questions with expected answers before building anything.
If you cannot answer the last one, you are not ready to build. Evaluation is not a final step; it is what tells you whether retrieval is helping at all.
Build the smallest version first
A workable first version is smaller than most teams expect: a few dozen documents, a simple chunking strategy, an off-the-shelf embedding model, and a prompt that forces the model to say "I don't know" when the retrieved context is insufficient. That last instruction matters more than any vector database choice.
Measure it against your test questions. If accuracy is poor, check the corpus and the chunking before reaching for a more complex architecture. Most RAG problems are data problems wearing a technical costume.
Let's talk about your case
If you are weighing RAG against a simpler approach — or you have a prototype that answers confidently and wrongly — the useful next step is a short conversation about your actual documents and questions. No pitch, just a look at whether retrieval is the right tool for what you are trying to do. You can reach me at contact@hamzabelgacem.com.