---
title: "Theoretical limits can expose inference problems"
date: 2026-08-09
canonical: https://solmaz.io/x/2086333655119720782/
x_url: https://x.com/onusoz/status/2086333655119720782
license: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
---

You can use the calculator in https://localmaxxing.com/en now to predict theoretical upper bounds from https://solmaz.io/llm-performance-upper-bounds 🙌

A little note, pay attention to the speculative decoding coefficient rho. If an upper bound is calculated using rho=1, then dspark might push it well above that limit. For example, the formulation predicts an upper bound of 28 decode tok/s, but dspark pushes it up to 35-50 tok/s in https://github.com/0xSero/deepseek-v4-flash-0731-spark-sparkinfer by @0xSero  

I am curious whether we will observe a phenomenological law between real-life engine performance and the upper bound, like real-life performance maxes out at e.g 80% of the upper bound in most cases

That would be very useful! Then we would be able to tell when something is wrong with an inference engine, if it doesn't reach at least 80% of the upper bound, ignoring spec. decoding

*Part 1/2 of a thread; root: https://solmaz.io/x/2086333655119720782/*
*Quotes a post by @LottoLabs (https://x.com/LottoLabs/status/2086237588952920130); its text is omitted here because it is not covered by this site's license.*
