Production RAG Architecture: Retrieval Pipelines That Actually Work
Ingestion, chunking, embeddings, hybrid search, reranking and context assembly, plus the tenancy leaks and injection paths each stage quietly opens up.
· 16 minDesk · Labs & reverse engineering
Technical deep dives, reproducible tests, and tool evaluations.
Ingestion, chunking, embeddings, hybrid search, reranking and context assembly, plus the tenancy leaks and injection paths each stage quietly opens up.
· 16 minA hosted assistant is a product, not a model. Here is which of its workloads open weights already match, which they still do not, and what control buys you.
· 14 minA workload-first reasoning guide to the 96 GB single-address-space card: what genuinely fits, where MoE models break it, and when a cluster or rental wins.
· 12 minA 135 GB model against a 128 GB box forces a cluster. Here is the decision order, the parallelism maths, and what the interconnect really costs per token.
· 13 minPaged KV, radix-tree prefix sharing and ahead-of-time compilation all pull in different directions. Work out which one your own workload should pay for.
· 14 minA model that fits can still die on one long prompt. Here is the arithmetic behind the second allocation that grows with context, and how to shrink it.
· 11 minWeights are only one of four memory buckets, and past roughly 115k tokens the KV cache costs more than the model does. Here is the arithmetic in full.
· 16 minPrefill is compute-bound, decode is bandwidth-bound, and KV cache is what really caps concurrency. The metrics, memory maths and tradeoffs behind serving.
· 16 min