Weekly AI & Engineering Digest — Sep 6, 2026
by Vamshi • 9/6/2026GPT-6 Astra's benchmarks and safety classification, attention-mechanism internals, embedding compression, RAG retrieval failures, how Shopify beat GPT-5.6 with a 0.8B model.
Read PostGPT-6 Astra's benchmarks and safety classification, attention-mechanism internals, embedding compression, RAG retrieval failures, how Shopify beat GPT-5.6 with a 0.8B model.
Read PostvLLM vs SGLang vs Ollama, GraphRAG map-reduce search, a one-line Qdrant fix for multivector memory, and why agent skills work as runbooks, not facts (8,100 trials).
Read PostA cheaper model can double your turn cost, vLLM's 23x batching trick, Google's split TPU 8t/8i chips, and the 3-layer stack teams are shipping to sandbox AI agents.
Read PostKV cache economics, the LLM security threat map, 5 models on one GPU, and a 7B model beating a 32B via distillation — six deep dives from Daily Dose and ByteByteGo this week.
Read PostHow ChatGPT's agent loop cuts token costs, 6 algorithms that auto-optimize LLM prompts, and how DoorDash, Instacart and Uber Eats each wired LLMs into search differently.
Read PostFive quantization techniques for 70B models on one GPU, the inference optimizations production stacks need, and building a custom RAG eval metric from scratch.
Read PostRebuilding Claude Code's 6-layer harness in CrewAI, why small models alone won't cut your inference bill, and the four agent loop patterns — plus Kimi K3's 2.8T drop.
Read PostKV caching rethought with LMCache and 14x faster TTFT, the 4-stage production routing pipeline, and Fable-as-Advisor at 92% quality for 63% of the cost.
Read PostModel routing that cuts costs 50-60%, the 4-layer agent engineering stack, and self-improving harnesses lifting small models 33-60% — plus Sonnet 5 ships.
Read PostSpeculative decoding at 4x speedups, the RAG taxonomy every engineer needs, and Karpathy's agentic engineering framework — plus GPT-5.6 under government review.
Read PostStructured notes from the week's best AI and engineering newsletters — retrieval layers as a shared tool, 8.5x faster LLM inference, and model distillation.
Read Post