If you have ever shipped a data pipeline, you already understand RAG. Not the hype version, the production version. It is ingest, chunk, embed, index, retrieve, generate, with an LLM bolted on at the end. The model is the last box in the diagram. Everything before it is your world.
That is the thesis this whole site is built on. Most RAG failures are data engineering failures wearing an AI costume. Bad chunking strategy, a stale vector index, garbage source data. No model upgrade fixes any of that. Retrieval quality caps answer quality. Full stop.
One pipeline, no magic step
Draw the architecture and it looks like every pipeline you have ever built. Documents, PDFs, and tickets go in one end. Something splits them into passages. Something turns passages into vectors. Something indexes those vectors. A query comes in, the index returns the top-k matches, and the model generates an answer from that context.
There is no step in there that requires a PhD. The LLM does the last five percent of the work, the part that turns retrieved passages into sentences. The other ninety-five percent is moving data from one box to another without corrupting it. You have been doing that for years.
You already do most of this
Translate the vocabulary and the mystery evaporates.
Chunking is picking your grain. Every warehouse model you have ever designed started with the same question: what is one row? A document chunk is the same decision. Too coarse and you smuggle noise into every retrieval. Too fine and you shred the context the model needs to answer.
Embeddings are feature engineering. You take raw text and turn it into numbers that preserve the relationships you care about. You have done this with every categorical variable you ever encoded. The vector is just a longer feature row.
The vector database is an index. It is an approximate nearest neighbor index, but operationally it is the same thing as every other index you maintain: a structure that makes a specific lookup fast, with freshness and quality characteristics you have to monitor.
Top-k retrieval is ORDER BY similarity LIMIT k. That is not a simplification. That is what the query is.
So when someone tells you RAG is an AI problem, ask which of the six stages they mean. Five of them live in your codebase already.
It breaks at chunking
Production failure number one, and it is not close. Split too big and every retrieved chunk drags in noise, which means the model answers from a haystack with your needle taped to the side. Split too small and the chunk has no context, so the model gets the right passage and still cannot use it. Either way the model looks stupid and the model was fine.
Your chunk grain decides what retrieval can ever find. Not what it does find on a good day. What it can find, ever, on any query, because the grain is baked into the index before the first question is asked. This is a schema design decision with the permanence of a schema design decision. Treat it like one.
Stale index, wrong answers
Production failure number two. The docs update. The index does not. The model answers today’s question from yesterday’s data, confidently, with citations to the old version. Nobody notices until a customer does.
This is pipeline lag wearing a new hat. You would never let a warehouse table go stale without an alert. The vector index is a derived table. It gets a freshness SLA and an alert like every other derived table, or it goes stale exactly when it matters most.
RAG cannot fix bad data
The honest limit, and the one that makes this post the account thesis instead of a tutorial. Garbage in, garbage out did not retire when the transformer arrived. If the source documents are wrong, contradictory, or missing, the best retrieval in the world returns wrong, contradictory, or missing context, and the model generates a confident answer from it.
No model outranks the context it is given. A bigger model does not fix a broken pipeline. It just fails more eloquently.
That is why the data engineer owns this system, not the prompt engineer. The quality bar is a data quality bar. The monitoring is pipeline monitoring. The failures are the failures you have seen before, in ETL jobs and warehouse loads, back when nobody called them AI.
We started @dataeng.ai for exactly this. AI for data engineers, no hype, just what works in production. This post is the map. Everything else on this site is the territory.
This post started as an Instagram post →
Numbers above trace to these sources. If one moved, tell us and we fix it.


Talk it through
Argue with us on Instagram.