◆ KEYNES SOFTWARE

The RAG Bill — Where the Money Goes in a Retrieval Pipeline

2026-07-07 · ◐ THRIFT-1 · coords [0.52, 0.28] · EN
> TRANSMISSION RECEIVED · PLANET THRIFT-1

TRANSMISSION RECEIVED · PLANET THRIFT-1 · COORDS [0.52, 0.28]

The previous planet showed how RAG lets a brand knowledge base answer open-book questions. Land on THRIFT-1 and the bill arrives fast — every query runs embedding, retrieval, and generation, each stage burning compute. The mission here is simple: see where the money actually goes, and where to cut without losing answer quality.

Where the cost lives in a RAG pipeline

A retrieval-augmented generation pipeline has three persistent cost centers. Embedding turns every product manual, brand guide, and FAQ into vectors at ingestion time — the larger the corpus and the larger the model, the more CPU and GPU it consumes. The vector store (Redis, Milvus, Pinecone) bills by scale and query volume, and wider recall (a larger k) costs more per query, while recall that is too narrow risks missing the answer entirely. The final LLM generation step is where the real money hides: a 70B-class model on a flagship accelerator is the electricity bill of the entire pipeline.

The HuggingFace and Intel benchmark delivers a counterintuitive result. In an enterprise RAG workload running Llama2-70B at 16 concurrent requests, an NVIDIA H100 reaches only 1.13x the throughput of a Gaudi 2 — but its performance per dollar is just 0.44x of Gaudi 2. In plain terms: the H100 costs more than twice as much per query while barely running faster. Hardware choice, not model size, is the single biggest cost lever in a RAG deployment, and ignoring it is the fastest way to inflate a knowledge-base budget.

The next lever is quantization. Dropping from BF16 to FP8 pushes Gaudi 2 throughput up another 1.8x — the same answer quality at roughly half the unit cost. Quantization is invisible to the application layer, which means the only real question is whether your inference provider exposes it. A vendor that silently runs BF16 because it is simpler is leaving a 2x cost reduction on the table.

The third lever is moving embedding off the GPU entirely. Embedding models are small and bursty — a product catalog refreshes once a week, not once a millisecond — so parking them on a flagship card is waste. Running embeddings on a CPU with the AMX instruction set (Granite Rapids) yields 2-3x performance gains for mixed AI workloads, leaving the expensive accelerators free for the generation step that actually needs them. None of these are research-grade tricks; they are procurement decisions: which accelerator, which precision, which silicon runs which stage. A pipeline that ignores them pays 2-4x more per query for identical answer quality — a tax that compounds with every product launch and every FAQ update.

Why marketers should care

Marketing teams rarely see the inference bill, but they feel it as capped pilots and “we can’t afford to index the full catalog” conversations. RAG cost structure decides how deep a brand knowledge base can go: cheap embedding means every SKU, every localized variant, every campaign footnote can live in the index; expensive generation means the team must ration queries or accept a duller model. The marketing-relevant metric is not API unit price but cost per thousand answered queries — and that number is set by hardware selection and quantization, not by prompt engineering. A vendor quoting GPU hourly rates without mentioning FP8 or CPU-side embedding is selling runway, not range.

How to use it

  • Split the bill three ways — embedding, vector store, generation — and negotiate each separately instead of accepting a bundled API price.
  • Ask vendors for “cost per 1,000 answered queries” and whether FP8 and CPU-side embedding (AMX) are enabled; put both in the SLA.
  • Default to a mid-tier accelerator (Gaudi 2-class) for generation and reserve flagship GPUs only when answer quality demonstrably drops.

// END OF LOG

// END OF LOGS