Series

RAG Architectures

Retrieval-augmented generation from ingestion through to the answer. The series treats chunking, embeddings, vector stores, reranking and evaluation as architecture questions, because they are what decides whether a knowledge base returns dependable answers or merely plausible ones.

12 articles

What makes up an AI assistant's response time What makes up an AI assistant's response time An AI assistant's response time is made up of five items. Measuring each one on its own finds faster fixes than new hardware or a different model. AI 8 min 09/14/2026 Why AI features fail in practice: seven stages where errors arise Why AI features fail in practice: seven stages where errors arise A wrong AI answer almost always has an address. Knowing the seven stages of an AI feature lets you find the fault before anyone swaps the model. Methodology 9 min 09/14/2026 Was ein RAG-System im Betrieb wirklich kostet Was ein RAG-System im Betrieb wirklich kostet Nicht das Modell ist der Kostenblock. Es sind Indexpflege, Evaluation und der Betrieb, und diese drei stehen in keinem Angebot, das nur die Entwicklung beziffert. Methodik 9 min 08/08/2026 Baidu OCR in the stack: why classic text recognition stays alongside visual search Baidu OCR in the stack: why classic text recognition stays alongside visual search OCR is not dead. PaddleOCR-VL delivers searchable text where visual search alone falls short. Tools 2 min 06/25/2026 F-RAG (RAG-Fusion): how it differs from plain RAG F-RAG (RAG-Fusion): how it differs from plain RAG RAG-Fusion runs several query variants and fuses the results with reciprocal rank fusion. Better recall, some drift risk. Methodology 2 min 06/25/2026 Late interaction explained: why Qdrant and ColQwen build the better knowledge base Late interaction explained: why Qdrant and ColQwen build the better knowledge base Late interaction compares every search term against every page excerpt. Qdrant stores that natively. Methodology 2 min 06/25/2026 The OCR-free document stack 2026: ColPali, ColQwen, ModernVBERT, and Qdrant The OCR-free document stack 2026: ColPali, ColQwen, ModernVBERT, and Qdrant Visual retrieval models search the page as an image instead of via OCR. A sober overview. Methodology 3 min 06/25/2026 Query decomposition and the advanced-RAG toolkit: match the method to the failure Query decomposition and the advanced-RAG toolkit: match the method to the failure Query decomposition, HyDE, RAG-Fusion, GraphRAG each fix one failure. The sophisticated 2026 RAG knows when to use none. Methodology 4 min 06/25/2026 The end of OCR? Visual document search and what it changes for the Mittelstand The end of OCR? Visual document search and what it changes for the Mittelstand Visual retrieval models find tables and scans where OCR-based RAG fails. What that means in practice. AI 2 min 06/25/2026 Subquadratic LLMs: cheaper long context for on-prem RAG Subquadratic LLMs: cheaper long context for on-prem RAG Subquadratic models lower the cost of long context. A real lever for on-prem RAG, but still young. Research 5 min 06/18/2026 What is RAG? Retrieval-augmented generation explained for the Mittelstand What is RAG? Retrieval-augmented generation explained for the Mittelstand RAG couples a language model to your own documents, so answers stay verifiable and current. AI 4 min 06/16/2026 Retrieval-Augmented Generation, and why it underpins sovereign AI Retrieval-Augmented Generation, and why it underpins sovereign AI RAG grounds a model's answers in retrieved sources, cutting hallucination and staleness. Why it underpins sovereign AI. AI 5 min 04/01/2024

Wayne Dyer

“If you change the way you look at things, the things you look at change.”