Speculative decoding
Mentioned in
- Theoretical limits can expose inference problems
- Benchmarking Laguna S 2.1 on DGX Spark
- Report Per-Session and Aggregate Token Throughput
- Upper Bounds for Local Model Throughput
- Write-up of Reiner Pope's Lecture: How GPT, Claude, and Gemini Are Actually Trained and Served
- Towards 1-click setup for local models in OpenClaw