Weekly AI & Engineering Digest — Jul 27, 2026
by Vamshi • 7/26/2026Five quantization techniques for 70B models on one GPU, the inference optimizations production stacks need, and building a custom RAG eval metric from scratch.
Read PostFive quantization techniques for 70B models on one GPU, the inference optimizations production stacks need, and building a custom RAG eval metric from scratch.
Read PostRebuilding Claude Code's 6-layer harness in CrewAI, why small models alone won't cut your inference bill, and the four agent loop patterns — plus Kimi K3's 2.8T drop.
Read PostKV caching rethought with LMCache and 14x faster TTFT, the 4-stage production routing pipeline, and Fable-as-Advisor at 92% quality for 63% of the cost.
Read PostModel routing that cuts costs 50-60%, the 4-layer agent engineering stack, and self-improving harnesses lifting small models 33-60% — plus Sonnet 5 ships.
Read PostSpeculative decoding at 4x speedups, the RAG taxonomy every engineer needs, and Karpathy's agentic engineering framework — plus GPT-5.6 under government review.
Read PostStructured notes from the week's best AI and engineering newsletters — retrieval layers as a shared tool, 8.5x faster LLM inference, and model distillation.
Read Post