Entries for July 14, 2026

  1. Portrait of Onur Solmaz
    Onur Solmaz · Post · /2026/07/14

    Write-up of Reiner Pope's Lecture: How GPT, Claude, and Gemini Are Actually Trained and Served

    Note: this post is an AI-assisted write-up of the blackboard lecture Reiner Pope gave on Dwarkesh Patel’s podcast. Watch the original video: How GPT, Claude, and Gemini are actually trained and served (YouTube, 2h13m).1

    Pope is the CEO of the chip startup MatX and previously worked on TPU architecture at Google. With two rules of thumb (a roofline model of a GPU rack, and “set competing costs equal to each other”), he derives why batching makes tokens up to 1000x cheaper, why frontier models may be over-trained ~100x beyond Chinchilla-optimal, and how much of a lab’s serving stack you can reverse-engineer from its public API prices. The figures below are redrawn from the blackboard.

    I have tried to stay faithful to the original throughout, converting the dialogue into prose and keeping all the numbers as stated. Any errors introduced in the conversion are mine.

    The question that motivates everything

    Dwarkesh opens with a pricing puzzle. Companies like Anthropic, OpenAI, and Cursor offer a “fast mode” that streams tokens at roughly 2.5x the speed for 6x the price. What is mechanically going on that makes this trade possible? Could you pay 100x more and go even faster? And could there be a “slow mode” where you wait minutes and pay much less?

    Pope’s answer is that the dominant effect is batch size, and the rest of the lecture quantifies exactly what batching does to latency and cost. (A second effect, speculative decoding / multi-token prediction, is set aside.)

    The whole analysis rests on two simplifications:

    1. A roofline model of the hardware. For a cluster like an NVIDIA Blackwell NVL72 rack (72 GPUs), only two numbers matter: memory bandwidth and compute throughput (FLOPs).
    2. Two numbers for the model. The time to operate on the weights, and the time to operate on the context (the KV cache).

    The KV cache is the per-conversation state the model keeps in memory. During decode, each new token runs a full forward pass through all the weight matrices, and its attention mechanism looks back at an internal representation of every previous token. That stored representation is the KV cache, and reading it is dominated by memory fetches rather than matrix multiplies.

    The two-line roofline

    The time for one decode step is bounded below by whichever is slower, the memory system or the compute:

    tmax(tmem,  tcompute)t \geq \max(t_{\mathrm{mem}},\; t_{\mathrm{compute}})

    The compute side has to multiply a batch of BB tokens by all the active parameters:

    tcompute=BNactiveFLOPst_{\mathrm{compute}} = \frac{B \cdot N_{\mathrm{active}}}{\mathrm{FLOPs}}

    (The attention compute is ignored; it is small in comparison.) Note the distinction between active and total parameters: in a mixture-of-experts model like DeepSeek V3, about 37B parameters are active per token out of roughly 700B total.

    The memory side has to fetch all the weights once per step, plus the KV cache of every sequence in the batch:

    tmemNtotal+Blenctxbytestokmemory bytes/st_{\mathrm{mem}} \geq \frac{N_{\mathrm{total}} + B \cdot \mathrm{len}_{\mathrm{ctx}} \cdot \mathrm{bytes}_{\mathrm{tok}}}{\text{memory bytes/s}}

    These two lines are enough to draw the latency picture:

    Decode latency vs batch size

    The weight fetch is a constant floor: no matter how small the batch, you must stream all total parameters from HBM into the chips once per token, and if you use all your memory bandwidth you cannot beat that. This is the latency lower bound, and it already answers the fast-mode question: for a given hardware configuration there is a floor on how fast tokens can come out, and paying more only helps until you hit it.

    Cost is a different plot. Renting the GPUs for one step costs the same regardless of batch size, but the step produces BB tokens, so the cost per token is t/Bt/B:

    Cost per token vs batch size

    At batch size 1 the weight fetches are not amortized over anything and the economics are up to a thousand times worse. As the batch grows, the weight-fetch hyperbola vanishes and the compute term becomes a hard cost floor. This also answers the “slow mode” question: a hypothetical Claude Code Slow would live on that floor, and it would not be much cheaper than normal serving, because the compute and the KV fetches are unique to each request and cannot be amortized further.

    The magic batch size

    Where is the balance point where memory time equals compute time? Ignoring the KV term for a clean answer and equating the weight fetch with the weight multiply:

    Ntotalmem BW=BNactiveFLOPsB=FLOPsmem BWNtotalNactive\frac{N_{\mathrm{total}}}{\text{mem BW}} = \frac{B \cdot N_{\mathrm{active}}}{\mathrm{FLOPs}} \quad\Longrightarrow\quad B = \frac{\mathrm{FLOPs}}{\text{mem BW}} \cdot \frac{N_{\mathrm{total}}}{N_{\mathrm{active}}}

    The first factor is purely a hardware constant. Counted in FP4 multiplies (half a byte each), it comes out around 300 on most GPUs, and it has stayed roughly stable from A100 to H100 to B100 because FLOPs and memory bandwidth grew together. The second factor is the sparsity of the model. So:

    B300×sparsityB \gtrsim 300 \times \mathrm{sparsity}

    For DeepSeek, which activates 32 of 256 experts (sparsity 8), that gives a batch of about 2,400 sequences. In practice people run double or triple that, since real-world efficiency is worse than the roofline. Including the KV fetch would push the optimal batch higher still. Remarkably, this result depends only on sparsity, never on model scale.

    Trains departing every 20 milliseconds

    How does a batch fill up with real users? Pope’s model is a train schedule. The server starts a new batch every ~20 ms whether or not it is full: any requests that are ready board the train, and a request that arrives just after departure waits for the next one. Worst-case queueing latency is therefore about 40 ms.

    The 20 ms itself comes from a separate design principle: you want to read your entire HBM capacity once per forward pass, so the natural step time is capacity divided by bandwidth. On the Rubin generation that is 288 GB / 20 TB/s ≈ 15 ms, and the number has hovered around 20 ms across many HBM generations. There is no point going slower, because reading the read-only weights or the KV cache twice per token does nothing for you.

    A batch of ~2,000 at ~64 steps per second is ~128,000 tokens per second per rack. Google has bragged about Gemini traffic in the hundreds of millions of tokens per second worldwide, so one rack’s economical batch is about one-thousandth of Gemini. That is the economy of scale in inference: real, but reachable by any serious provider.

    Does sparsity hurt quality?

    The roofline says sparsity is nearly free performance, so the follow-up is empirical: how much quality do you lose? From the paper “Unified Scaling Laws for Routed Language Models”, with an older MoE technique, a 64-expert model with 370M active parameters matched a dense 1.3B model. That is a 64x increase in total parameters for a 4x effective gain, a huge parameter cost for a modest efficiency win.

    And yet from the systems side it is still nearly a pure win: the extra weight fetches amortize over a larger batch, so you keep increasing sparsity until you run out of simultaneous users. The real price is memory capacity, which is what the next sections are about.

    Laying out a mixture of experts on a rack

    An MoE layer has a router that sends each token to a small fraction of the experts (each expert being an ordinary MLP), an all-to-all “dispatch” of tokens to their experts, an all-to-all “combine” that sums the results, and a residual connection around the whole thing.

    The standard practice is expert parallelism: different experts live on different GPUs. DeepSeek’s 256 experts on a Blackwell rack (using 64 of the 72 GPUs for divisibility) means 4 experts per GPU. Since the router’s decisions are data-dependent, any GPU may need to send tokens to any other GPU.

    MoE layer with expert parallelism

    This all-to-all traffic pattern is a perfect fit for how a rack is wired. In NVIDIA’s design the GPUs sit on the outside of the rack and NVSwitches in the middle, with every GPU cabled to every switch, so any GPU reaches any other in two hops. This is the scale-up network (NVLink). Leaving the rack means taking the scale-out network through a NIC and a data-center switch, which is typically about 8x slower.

    Scale-up vs scale-out

    If you spread one expert layer across two racks, half of every all-to-all crosses the slow rack-to-rack boundary and becomes the bottleneck. So one rack bounds the size of an expert layer, and this is what has been driving interconnect domains bigger: Hopper had 8 GPUs in a scale-up domain, Blackwell 72, Rubin 500-something (some of that is Jensen math, but there is a genuine ~4x from a much harder rack design). The physical constraint is mundane: cable density. Doubling the GPUs in a rack literally doubles the density of cables that must be routed to the switches, against limits of space, weight, power, cooling, and the bend radius of the cables.

    This is also a lens on model scaling history. GPT-4 (2023) was rumored to be over a trillion parameters, and models only clearly exceeded that scale once racks with tens of terabytes of fast memory arrived. Google’s TPU deployments have had very large scale-up domains for a long time, which may be part of why Gemini’s pre-training scaled successfully early. The summary: active parameters are limited by compute cost, and total parameters are limited by scale-up size.

    Pipeline parallelism

    Expert parallelism uses up one rack. To use more racks, the remaining options are data parallelism and pipeline parallelism (tensor parallelism has become irrelevant now that experts are small). Pipelining means putting different layers on different racks: a token flows through rack 0 for the first stage of layers, then hops to rack 1, and so on.

    Is the hop a bottleneck? Compare the time spent on scale-up traffic to the time on scale-out traffic. Crossing racks sends each token once per stage, while inside the rack each token fans out to every activated expert, twice (dispatch and combine), for every layer in the stage:

    tscale-uptscale-out=182(activated experts)(layers per stage)    1\frac{t_{\text{scale-up}}}{t_{\text{scale-out}}} = \frac{1}{8} \cdot 2 \cdot (\text{activated experts}) \cdot (\text{layers per stage}) \;\geq\; 1

    The 1/8 is the bandwidth ratio. With 8+ activated experts and a few layers per stage, the inequality is easily satisfied, so an entire pipeline of racks, one stage each, is communication-feasible.

    Dwarkesh brings up Ilya’s remark that “as we now know, pipelining is not wise,” and the architectural constraints it imposes (e.g. Kimi’s attention to layers a few back is awkward to pipeline). Pope’s framing: pipelining is a massive hassle with real but narrow benefits. It saves no runtime at all (the memory fetches just happen on a different rack), but it divides the weight storage per rack, which matters if memory capacity is your constraint.

    The catch is micro-batching. To keep four racks busy, you need four micro-batches in flight, each wrapping around for its next decode step as soon as it finishes:

    Pipeline schedule during decode

    In inference this is natural and the bubble costs nothing; latency is identical to running unpipelined on one rack. In training there is a hard stop between the forward and backward passes of a batch, which creates a genuine bubble of idle time (the literature has zero-bubble and one-forward-one-backward schemes to interleave around it; as Dwarkesh notes, you could also mine Bitcoin in it):

    Pipeline schedule during training

    The training batch size itself is a trade-off: smaller batches are always better for ML convergence (fresher gradients), larger batches are better for systems throughput, and the optimum sits in between.

    The memory wall and why the KV cache won’t shard

    Here Dwarkesh raises the macro puzzle. Memory is the scarce commodity of the moment: Dylan Patel claims hyperscalers are spending half of their CapEx on memory, and consumer devices are getting squeezed. Yet the pipelining analysis just said racks have a memory surplus, since a trillion-parameter model needs only ~1 TB against a rack’s tens of TB. Why is Jensen shoving all that HBM in?

    Write down the memory demand across the whole system:

    Cmem=Ntotal+BlenctxbytestokC_{\mathrm{mem}} = N_{\mathrm{total}} + B \cdot \mathrm{len}_{\mathrm{ctx}} \cdot \mathrm{bytes}_{\mathrm{tok}}

    Sharding across EE GPUs of expert parallelism and PP racks of pipelining, the per-GPU requirement is this divided by EPE \cdot P. But the global batch is (number of micro-batches) × (micro-batch size), and the number of micro-batches needed to fill the pipeline equals PP, while the micro-batch size bb is pinned near 300×sparsity300 \times \mathrm{sparsity} by the roofline. Substituting B=PbB = P \cdot b, the PP‘s cancel in the KV term:

    cmemper-GPU=NtotalEP+blenctxbytestokEc_{\mathrm{mem}}^{\text{per-GPU}} = \frac{N_{\mathrm{total}}}{E \cdot P} + \frac{b \cdot \mathrm{len}_{\mathrm{ctx}} \cdot \mathrm{bytes}_{\mathrm{tok}}}{E}

    More pipeline stages keep shrinking the weight footprint, but the KV footprint per GPU stays constant: each extra stage requires proportionally more sequences in flight to stay busy. The KV cache can’t be amortized across the batch (it is unique per user), and it can’t be sharded across pipeline stages either. It loses on both fronts, and once you pipeline even a little, it becomes the dominant use of memory.

    So what do labs actually run? Per the DeepSeek paper: expert parallelism up to the scale-up domain size, then very little pipelining (maybe none, maybe 2 stages so the weights aren’t an issue). Frontier inference essentially lives inside a single scale-up domain. Each rack hop would also add on the order of milliseconds of latency per token, which stacks across stages in sequential decode.

    The last piece of the scale-up story is bandwidth. The weight-fetch latency is

    tmem, weights=NtotalS×BW per GPUt_{\text{mem, weights}} = \frac{N_{\mathrm{total}}}{S \times \text{BW per GPU}}

    where SS is the scale-up size, because all GPUs in the domain load the weights in parallel. Per-GPU HBM bandwidth improves maybe 1.5–2x per generation, but SS jumped 8x from Hopper to Blackwell. Pipelining solves the capacity problem; big scale-up domains solve the bandwidth problem, which is what actually lets you serve at low latency and long context.

    Over-trained 100x beyond Chinchilla

    Chinchilla scaling tells you the compute-optimal ratio of model size to training data. But a lab does not minimize training compute; it minimizes total compute across pre-training, RL, and inference for all its users. Pope’s heuristic: when minimizing a sum of competing costs, the minimum tends to sit where the costs are equal (true for x+1/xx + 1/x, for ex+exe^x + e^{-x}, and generally for power laws). So set all three equal.

    Using the 6ND rule (6 FLOPs per parameter per token for forward+backward, 2 for forward only):

    ctotal=6NactDPTpre-training  +  [2 to 6]NactDRLinefficiencyRL  +  2NactDinfinferencec_{\mathrm{total}} = \underbrace{6\, N_{\mathrm{act}} D_{\mathrm{PT}}}_{\text{pre-training}} \;+\; \underbrace{[2\text{ to }6]\, N_{\mathrm{act}} D_{\mathrm{RL}} \cdot \text{inefficiency}}_{\text{RL}} \;+\; \underbrace{2\, N_{\mathrm{act}} D_{\mathrm{inf}}}_{\text{inference}}

    RL sits between 2 and 6 because you generate every rollout but may not train on all of it, and it carries an extra inefficiency factor (~30%) because RL involves a lot of decode, which runs at lower MFU than training. The active parameter count divides out entirely. Working through the arithmetic on the board, the equal-cost condition lands at roughly

    DPT1.5DRLDinfD_{\mathrm{PT}} \approx 1.5\, D_{\mathrm{RL}} \approx D_{\mathrm{inf}}

    In words: the number of pre-training tokens, RL tokens, and lifetime inference tokens should all be about the same, within factors the analysis can’t resolve. (Dwarkesh’s gloss: every model should stream out roughly the sum of human knowledge that was streamed into it.)

    Now plug in real-world guesses. Global traffic of ~500M tokens/s, cut by 5–10x for one specific model in a family, gives ~50M tokens/s; times a two-month deployment life, that is roughly 2.6×10142.6 \times 10^{14}, call it 200T inference tokens. The rumor mill says frontier pre-training is ~150T tokens, which matches. With ~100B active parameters, Chinchilla would prescribe only ~2T tokens. The ratio is about 100x over-trained, derived almost from first principles. As Pope puts it, approximate everywhere, set A equal to B, and it’s kind of empowering how far that gets you.

    One asymmetry he flags: if your model might miss the frontier and get thrown away, the expected inference tokens shrink, so you should derate the inference term and err toward less over-training.

    Reading the infrastructure off API prices

    Since providers are incentivized to price close to cost (otherwise someone scoops them), public price sheets leak infrastructure details.

    The 200k context surcharge

    Gemini 3.1 charges 50% more per token beyond 200k context. Redraw the roofline as a function of context length at a fixed large batch: compute cost is flat (the attention FLOPs slope is negligible until millions of tokens), while memory cost grows linearly with the KV cache. The provider wants to be profitable at every context length, so a two-tier price is laid over a kinked cost curve, and the price bump should sit near the crossover where the model flips from compute-bound to memory-bound:

    Cost vs context length with two-tier pricing

    Assuming the crossover is at 200k, you can solve for the model’s KV bytes per token. Setting KV fetch time equal to compute time and cancelling the batch size:

    bytestok=mem BWFLOPsNactlenctx=1300100B200k1.7 kB per token\mathrm{bytes}_{\mathrm{tok}} = \frac{\text{mem BW}}{\mathrm{FLOPs}} \cdot \frac{N_{\mathrm{act}}}{\mathrm{len}_{\mathrm{ctx}}} = \frac{1}{300} \cdot \frac{100\mathrm{B}}{200\mathrm{k}} \approx 1.7\ \text{kB per token}

    Is ~2 kB/token plausible? The KV size is (number of unique attention contexts) × 2 × dheadd_{\mathrm{head}} × (KV heads). With dhead=128d_{\mathrm{head}} = 128 and 8 KV heads and a single global context shared across all layers (the Character AI trick, also used in Gemma), you get exactly 2 kB. Sparse attention gets there with bigger raw numbers divided by the sparsity. So the pricing page is consistent with a real architecture, if maybe a little on the small side.

    Input vs output prices

    Output tokens cost 3–5x more than input tokens. The two phases differ in tokens per forward pass: decode processes one new token per pass, prefill processes the whole prompt in one pass. Dividing the same roofline by the tokens per pass (len_pass) gives the cost per token: the compute term is flat, and the memory term is a hyperbola that only bites when len_pass is small.

    Prefill vs decode cost per token

    Prefill is compute-bound, decode is memory-bandwidth-bound, and a 5x price gap says decode at the provider’s operating point is deeply memory-bound: they are paying ~5x more per output token in memory time than the compute floor.

    This also explains the context length plateau. Contexts jumped from ~8k (GPT-3 era) to 100–200k around GPT-4 and have hovered there for a year or two, which suggests that is the balanced cost point. The barrier to 100M-token contexts (the “in-context learning is enough for AGI” scenario) is memory bandwidth and capacity, and HBM is not getting hugely better. Sparse attention (DeepSeek published one mechanism that effectively puts a square root on the KV term) is a big one-time improvement, but going too sparse costs quality, so it is a get-out, not a solution.

    Cache pricing and memory tiers

    Providers charge much less for cached input tokens, and charge different rates for keeping a cache alive 5 minutes vs 1 hour. There are two ways to produce a KV cache for a token: rematerialize it from scratch (a forward pass: pure compute cost, nothing to store) or hold it in some memory tier (near-zero retrieval cost, but you occupy capacity that scales with hold time). Each tier down (HBM → host DDR → flash → spinning disk) is cheaper to occupy and slower to retrieve from.

    Which tier backs which price? Pope’s rule: a storage tier is well-matched to hold times around its drain time, capacity divided by bandwidth, the same ratio that gave 20 ms for HBM. DDR drains in seconds, flash in about a minute, spinning disk in about an hour. So a 5-minute cache tier and a 1-hour cache tier probably map to flash and spinning disk, which surprised him: “I’m kind of shocked to see spinning disk being used at all.”

    Convergent evolution with cryptography

    The sit-down portion covers Pope’s blog post on how neural nets and ciphers evolved similar shapes. Both need to thoroughly mix information across all their inputs, and even stirring cake batter alternates directions for the same reason. But they optimize in opposite directions. A neural net is kept differentiable in a useful way: residual connections and LayerNorm exist to keep the derivative simple and meaningful for gradient descent. A cipher is designed so that its derivative is useless: differential cryptanalysis attacks a cipher by differentiating it (over the field of two elements), and a well-designed cipher makes a small input difference blow up into a huge output difference, the avalanche effect. Adversarial examples in image models are exactly the avalanche property showing up where it is not wanted.

    Building ciphers out of neural nets is a bad idea (99% of new ciphers get broken), but one construction has productively flowed the other way. A Feistel network turns any non-invertible function ff into an invertible two-input block:

    g(x,y)=(y+f(x),  x)g(x, y) = (\,y + f(x),\; x\,)

    To invert, read off xx from the second slot, then recover y=zf(x)y = z - f(x).

    Feistel construction

    The 2017 RevNets paper imported this into deep learning: make each layer a Feistel block (which turns out to look like a residual connection from two layers back) and the whole network becomes invertible. Training normally has to write every layer’s activations to HBM on the forward pass so that the backward pass can read them, a memory footprint linear in depth and often the largest one in training. An invertible network stores none of it: during the backward pass it undoes the forward pass in lockstep, rematerializing activations as needed.

    That is spending compute to save memory, the exact mirror image of the KV cache, which spends memory to save compute. Given where hardware is, the KV cache direction is usually the profitable one, which is a fitting last word for a lecture that is mostly about the price of memory.

    1. How this post was made, in the interest of transparency. I first had one agent (Codex) prepare the raw material. My prompt, verbatim:

      https://www.youtube.com/watch?v=xmkSf5IS-zw

      download using yt dlp. create a jpeg screenshot every 30 seconds and save it

      also get the transcription for it from dwarkesh’s website if it exists. if not, you can transcribe using the whisper model in bob@isengard

      save it in a folder. I want to prepare it for an another agent to prepare it for further processing

      That produced the video, 268 frames at 30-second intervals, and Dwarkesh’s published transcript (no Whisper needed). I then handed the folder to a second agent (Fable, in Cursor), which read the transcript, inspected the frames, and redrew the blackboard diagrams as matplotlib figures. My prompt, verbatim:

      convert this to a write-up that is faithful to the original. use the video and images if needed. create figures based on what is on the screen and such

      The draft first lived as a standalone GitHub document (“just make it standalone, dont add it to my blog”, then “create doc in ~/scratch repo, github markdown. i will read that. make sure it renders nicely”) before I asked for this post (“ok I want you to create a blog post citing the dwarkesh podcast in my blog, linking the youtube and just saying that this is a write-up of the video”). In between I reviewed the output and asked for fixes, for example on the cost figure (“the graph after this: the lines are a bit tight. could we make it clearer?”) and on figure rendering (“also, in the svgs, the space between some things are too much. if you can improve the text rendering in the svgs, it would be great”).

  2. Portrait of Onur Solmaz
    Onur Solmaz · Post · /2026/07/14

    Theoretical Upper Bounds for LLM Performance

    Note: this post is AI generated, adapted from my working notes behind the Local Frontier calculator. It is a work in progress and may lack rigor in places. If you find an issue in the formulation, please email me at [email protected].

    How fast can a machine serve a large language model? I built Local Frontier, an interactive database of local AI hardware and model profiles, to answer that question for any hardware and model pair. This post derives, from first principles, the math the calculator runs. That math is a roofline, an upper bound that says what a given hardware and model pair cannot exceed under a stated set of assumptions. Nothing in it is specific to local machines. The bounds are built from memory capacity and bandwidth alone, so the same formulas cover a MacBook and a datacenter GPU node. The post focuses on local hardware because that is what Local Frontier compares. Its outputs are ceilings for comparing machines, and a real implementation lands somewhere below them.

    The argument builds in layers, each created by a problem the previous layer cannot solve. A memory system is two numbers, capacity and bandwidth, and their product turns out to be a natural measure of what the system can fund. That product yields a clean but loose throughput ceiling for batched decoding. The loose ceiling ignores per-session context traffic, so it is too generous for real serving, and repairing it gives the bound the calculator actually uses. The repaired bound is then filled in per architecture by small adapters and checked against real hardware.

    Interactive figures accompany the derivation. They all use the same toy setup, a 32B-class dense model at 4-bit (18 GB of weights, 6 GB of runtime overhead, 0.26 GB of KV per thousand tokens of context) on a 128 GB machine with 800 GB/s of sustained bandwidth, so the numbers stay comparable from figure to figure.

    Three questions

    Serving an LLM raises three separate questions, and mixing them up is an easy way to be wrong about a machine.

    QuestionWhat it asks
    Resident fitCan the model plus runtime overhead be held in memory at all?
    Single-session speedWhat is the memory-side ceiling for one active conversation?
    Useful serving throughputAcross many active sessions, how many tokens per second can the device produce while each session stays above a minimum useful rate?

    The third question is the hard one. A machine can fit many sessions in memory and still be too slow per session at that concurrency. So the calculator must report both whether sessions fit and whether the fitting sessions are fast enough to be worth running.

    Every number produced here is an upper bound. A real system can fall below it for reasons the memory model deliberately ignores, among them compute limits, kernel quality, quantization overhead, scheduling, CPU involvement, paging, interconnects, and thermal throttling. The value of a clean upper bound is that it tells you the best case you are allowed to hope for, and therefore how much room an implementation still has.

    Memory power

    Before any model-specific detail, ask what a memory system fundamentally offers. When people compare accelerators for local inference they list many specs, from capacity and bandwidth through compute throughput, cache hierarchy, PCIe lanes, and thermals. For the decode phase of autoregressive generation, two of these recur in almost every bound. How much state can the memory hold, and how fast can it move that state? We start there.

    Two numbers and their product

    Model an idealized memory system by exactly two quantities.

    C=usable memory capacity,R=sustained memory bandwidth.C = \text{usable memory capacity}, \qquad R = \text{sustained memory bandwidth}.

    Capacity is a stock, the amount of resident state that can exist at once. Bandwidth is a flow, the amount of state that can be moved per second. They have different units and answer different questions, so any single product of them needs justifying before we rely on it.

    Define memory power as

    D=CR.D = C R .

    If CC is in GB and RR is in GB/s, then DD is in GB2/s\mathrm{GB}^2/\mathrm{s}. Power is meant in the colloquial sense of capability, as in computing power, and no watts appear anywhere in this model. The units look strange, so the rest of this section explains why this particular combination of CC and RR is the right scalar.

    For a feel of the scale, a 24 GB GPU at 1000 GB/s has D=24×1000=24,000 GB2/sD = 24 \times 1000 = 24{,}000\ \mathrm{GB}^2/\mathrm{s}. One thing to be careful with throughout is that every memory quantity in a given calculation must use the same unit system. The equations are identical in bytes, GB, or bits, and only the numeric value of DD changes.

    Feasibility theorem

    Here is the toy problem that justifies CRCR. The models in this formulation are deliberately simple, and they are meant as rules of thumb, useful when deciding on hardware for a model or when checking whether an existing setup is getting the most out of the hardware it runs on.

    A pure memory workload is a pair w=(h,r)w = (h, r), where hh is the resident information it must keep alive and rr is the information flow rate it must sustain. The workload is feasible on the system M(C,R)M(C, R) when the memory can both hold the state and carry the traffic. There are exactly two ways to fail.

    The first is a capacity failure. If a workload needs more resident state than the memory can store, it cannot fit, and no scheduling trick repairs that.

    hC.h \le C .

    The second is a flux failure. If a workload needs more traffic per second than the interface can deliver, it cannot be sustained, and no surplus capacity repairs that.

    rR.r \le R .

    In the idealized model these two conditions are also sufficient. If hCh \le C and rRr \le R, allocate hh of state and stream at rate rr. Therefore the feasible set is precisely the rectangle

    F(C,R)={(h,r):0hC, 0rR},F(C, R) = \{(h, r) : 0 \le h \le C,\ 0 \le r \le R\},

    whose area is

    μ(F)=CR.\mu(F) = C R .

    So memory power has a concrete meaning. It is the measure of the feasible workload region. A workload lives at a point (h,r)(h, r) in the plane, the system can serve every workload inside its rectangle and none outside, and the size of that rectangle is CRCR.

    A caution before building on this. The rectangle is a toy model, and the world is messier in both directions. Real hardware delivers only a fraction of its catalog bandwidth. Even a perfectly sequential read falls short of the spec number, and the achieved fraction depends on the access pattern, on whether enough parallel work is in flight to hide memory latency, and on the hardware itself. The sufficiency claim is idealized too, because the rate a machine reaches depends on which bytes are read and in what order. The bounds below survive this, since real inefficiency only pushes a machine further below its ceiling. The cost lands on comparisons between machines. If one machine sustains 85% of its spec bandwidth and another 60%, ceilings computed from spec numbers make the second machine look better than it really is. Read RR as sustained bandwidth where a measured number exists, and read every comparison in this post as approximate.

    A caveat on the metric

    One misreading is worth heading off. The product CRCR measures the workload-feasibility region, the set of jobs the device can support at an instant. A different question, “how many distinct memory histories can this device produce over a time TT”, has a different answer, namely the resident information plus the information streamed over that window,

    C+RT.C + R T .

    The feasibility-region reading is the one an inference-throughput bound will need, because serving means keeping sessions resident in memory and streaming data for them.

    A first decode bound

    The feasibility theorem describes static workloads. Decoding is dynamic, emitting tokens over time. This section turns memory power into a throughput ceiling for batched decoding, and in doing so shows where the product CRCR enters a real performance bound.

    A toy decoder

    Consider the decode phase of an autoregressive transformer, simplified to the memory system alone. Let

    W=model weight footprint,K=per-session KV/state footprint,W = \text{model weight footprint}, \qquad K = \text{per-session KV/state footprint},

    let bb be the batch size (number of concurrent sequences), and let ss be the number of decode steps per second. Each decode step emits one token per active sequence, so aggregate output throughput is

    T=bs.T = b \, s .

    Throughput is a product of two things, how many sequences run in parallel and how fast the shared model can be swept, and the memory system caps each factor separately.

    Capacity limit

    The model weights must be resident once, and every active sequence needs its own KV cache, the per-conversation attention state that grows with context length. So the resident state is W+bKW + bK, and it must fit.

    W+bKCbCWK.W + bK \le C \quad\Longrightarrow\quad b \le \frac{C - W}{K} .

    This is the capacity limit, and it sets the maximum parallelism. It is exactly the hCh \le C condition of the feasibility theorem, with h=W+bKh = W + bK.

    Bandwidth limit

    In a dense transformer, each decode step must apply the model weights. In the memory-bound idealization, applying the weights means streaming roughly WW of data per step. With bandwidth RR, the step rate obeys

    sWRsRW.s W \le R \quad\Longrightarrow\quad s \le \frac{R}{W} .

    This is the bandwidth limit, and it sets the maximum step rate. It is the rRr \le R condition, with the per-step traffic playing the role of the flux.

    Memory-power decode bound

    Multiply the two caps. Throughput is parallelism times step rate, and each is separately bounded, so

    T=bsCWKRW=R(CW)KW.T = b\, s \le \frac{C - W}{K}\cdot\frac{R}{W} = \frac{R\,(C - W)}{K W} .

    Substituting D=CRD = CR and factoring out CC exposes the memory-power term.

    TDKW(1WC).T \le \frac{D}{K W}\left(1 - \frac{W}{C}\right).

    When the model is much smaller than memory, WCW \ll C, the correction vanishes and the bound collapses to the memorable form

    TDKW.T \lesssim \frac{D}{K W} .

    This is the memory-power decode bound. The product CRCR appears because maximum throughput genuinely factors into maximum parallelism times maximum step rate.

    CKhow many sessions resident×RWhow many sweeps per second=CRKW=DKW.\underbrace{\frac{C}{K}}_{\text{how many sessions resident}}\times \underbrace{\frac{R}{W}}_{\text{how many sweeps per second}} = \frac{CR}{KW} = \frac{D}{KW} .

    The numerator D=CRD = CR is the machine. The denominator KWKW is the workload, the per-session state times the model sweep size. The bound reads as throughput is memory power divided by memory cost per active model-token.

    Scope of the bound

    The memory-power decode bound governs batched throughput. Set b=1b = 1 and it degenerates to

    TRW,T \lesssim \frac{R}{W},

    so for a single session only bandwidth matters and capacity is merely a fit constraint. That is the correct behavior, and it makes clear what each metric is for.

    Use caseThe metric that governs it
    Single-user local chatbandwidth RR, with capacity CC as a fit gate
    Largest model that fitscapacity CC first, then bandwidth
    Maximum batched decode throughputmemory power D=CRD = CR

    Memory power is the right scalar precisely for the third row. The next section explains why even that bound is too optimistic.

    Missing traffic

    The memory-power decode bound assumes the only per-token memory traffic worth counting is the model sweep WW, shared across the batch. Real decoding also reads each session’s growing KV cache, and that traffic is private, so it does not amortize over the batch. Ignoring it makes the bound promise throughput that long-context serving can never reach. This section introduces the correct per-token accounting and shows the memory-power bound falls out of it as a loose corollary.

    Universal resource bound

    Step back to the most general statement, which holds regardless of architecture. If each output token costs at least qminq_{\min} of unavoidable memory traffic and at least amina_{\min} of unavoidable compute, and the device delivers at most RR bytes/s and FF FLOP/s, then over all setups that fit,

    Tmaxmaxsetup fitsmin ⁣(Rqmin, Famin).T_{\max} \le \max_{\text{setup fits}} \min\!\left(\frac{R}{q_{\min}},\ \frac{F}{a_{\min}}\right).

    Two rooflines, and the workload lives under the lower of them. Throughout this post I assume decode is memory-bound, which is the usual case for local serving, and keep the memory roofline,

    TmaxRq,T_{\max} \le \frac{R}{q},

    where qq is the memory traffic per output token. Everything now reduces to estimating qq honestly. The focus on the memory side is also practical. A hardware catalog can collect capacity and bandwidth consistently across consumer and workstation devices, while comparable sustained-compute numbers are much harder to obtain.

    Bytes per token

    Split the per-token traffic into the two kinds that behave differently under batching. The model (or active-expert) weights are shared, since one sweep serves the whole batch, so their per-token cost is divided by the batch. The KV/context read is private, since each session reads its own cache, so its per-token cost is not divided at all. With WactiveW_{\mathrm{active}} the shared weight traffic per iteration, ρ\rho the tokens emitted per session per iteration (one for ordinary decoding), and Kread(L)K_{\mathrm{read}}(L) the private context traffic per output token at active context length LL,

    qsimple(b)=Wactivebρ+Kread(L).q_{\text{simple}}(b) = \frac{W_{\mathrm{active}}}{b\,\rho} + K_{\mathrm{read}}(L).

    The first term shrinks as the batch grows, because more sessions share each weight sweep. The second term does not move. A larger batch does not make any session’s context cheaper to read, and this asymmetry is the main reason long-context serving behaves differently from short-context serving.

    Interactive figure: bytes per output token as the batch and the read context change. Enable JavaScript to explore it.

    Memory-power bound as a corollary

    Treat qsimpleq_{\text{simple}} as an optimistic lower estimate of the true bytes per token, qsimple(b)qactual(b)q_{\text{simple}}(b) \le q_{\text{actual}}(b). Dividing RR by a smaller denominator gives a larger quotient, so substituting qsimpleq_{\text{simple}} keeps the result an upper bound.

    Tmaxmaxb:Wresident+bKstore+OC RWactivebρ+Kread(L).T_{\max} \le \max_{b\,:\,W_{\mathrm{resident}} + b K_{\mathrm{store}} + O \le C}\ \frac{R}{\dfrac{W_{\mathrm{active}}}{b\rho} + K_{\mathrm{read}}(L)} .

    Here the maximization runs over batches that fit in memory, with WresidentW_{\mathrm{resident}} the resident model footprint, KstoreK_{\mathrm{store}} the KV memory stored per session, and OO the runtime overhead. This is the honest simple bound, bandwidth divided by shared-per-token cost plus private-per-token cost, maximized over fitting batches.

    Now recover the previous section’s bound by deliberately throwing information away. Since Kread(L)0K_{\mathrm{read}}(L) \ge 0, dropping it only loosens the denominator.

    RWactivebρ+Kread(L)RWactivebρ=bρRWactive.\frac{R}{\dfrac{W_{\mathrm{active}}}{b\rho} + K_{\mathrm{read}}(L)} \le \frac{R}{\dfrac{W_{\mathrm{active}}}{b\rho}} = \frac{b\,\rho\,R}{W_{\mathrm{active}}} .

    The capacity constraint caps the batch at b(CWresidentO)/Kstoreb \le (C - W_{\mathrm{resident}} - O)/K_{\mathrm{store}}, so

    TmaxρR(CWresidentO)KstoreWactive=ρDKstoreWactive(1Wresident+OC).T_{\max} \le \rho\,\frac{R\,(C - W_{\mathrm{resident}} - O)}{K_{\mathrm{store}}\,W_{\mathrm{active}}} = \rho\,\frac{D}{K_{\mathrm{store}}\,W_{\mathrm{active}}}\left(1 - \frac{W_{\mathrm{resident}} + O}{C}\right).

    This is the memory-power decode bound again. It is what you get by discarding the private context term and using the largest batch that fits. That is why it can sit far above achievable throughput while remaining a true ceiling. The ordering is

    actual throughput    simple (KV-aware) bound    memory-power bound.\text{actual throughput} \;\le\; \text{simple (KV-aware) bound} \;\le\; \text{memory-power bound}.

    The memory-power bound is the orientation line. The KV-aware bound is the one to serve from, and the next section develops it into the operational calculator.

    KV-aware bound

    This section turns the simple bytes-per-token bound into the model the calculator runs. Three refinements are needed. Context must be split into the part that controls memory and the part that controls speed. The shared weight traffic must be allowed to grow with the batch, which matters for mixture-of-experts (MoE) models, models that route each token through a small subset of their weights. And the batch must be filtered so that we never count concurrency at which every session has become uselessly slow.

    Two context lengths

    A serving system usually reserves KV space for a long maximum context but reads, on average, a shorter active context. These two lengths drive different parts of the bound, so we keep them separate.

    Lalloc=reserved (maximum) context,Lread=average active context,LreadLalloc.\begin{aligned} L_{\mathrm{alloc}} &= \text{reserved (maximum) context}, \\ L_{\mathrm{read}} &= \text{average active context}, \\ L_{\mathrm{read}} &\le L_{\mathrm{alloc}}. \end{aligned}

    The allocation length controls how much KV memory each session reserves, and therefore how many sessions fit. The read length controls how much context each output token must stream, and therefore per-token cost. Collapsing them into one number either overcharges memory or overcharges speed.

    Model quantities

    A model contributes five quantities. Two are about fitting, two are about speed, and one is about decoding style.

    SymbolRole
    WresidentW_{\mathrm{resident}}Full resident footprint (for MoE, all resident weights, including the inactive experts)
    Wbatch(b)W_{\mathrm{batch}}(b)Shared weight traffic per decode iteration at batch bb
    Kalloc(Lalloc)K_{\mathrm{alloc}}(L_{\mathrm{alloc}})KV/cache memory reserved per session, which controls concurrency
    Kread(Lread)K_{\mathrm{read}}(L_{\mathrm{read}})Private context traffic per output token, which controls decode cost
    ρ\rhoTokens emitted per session per iteration (ρ=1\rho = 1 ordinary, ρ>1\rho > 1 speculative)

    The split between KallocK_{\mathrm{alloc}} and KreadK_{\mathrm{read}} mirrors the split between the two context lengths. Allocation controls how many sessions fit, read controls how fast each one decodes. Note that WbatchW_{\mathrm{batch}} now depends on the batch size bb, for a reason the adapter section explains.

    Memory-fit batch

    The first gate is whether sessions fit. Load the model, reserve overhead, and divide the remainder by the per-session allocation.

    bmem(Lalloc)=CWresidentOKalloc(Lalloc).b_{\mathrm{mem}}(L_{\mathrm{alloc}}) = \left\lfloor \frac{C - W_{\mathrm{resident}} - O}{K_{\mathrm{alloc}}(L_{\mathrm{alloc}})} \right\rfloor .

    This is memory-fit concurrency only. It is necessary but not sufficient, and it is exactly the trap that makes a machine look like it can serve a hundred sessions when it cannot serve them usefully.

    Interactive figure: how many sessions fit in memory as capacity and reserved context change. Enable JavaScript to explore it.

    Aggregate and per-session ceilings

    The per-token traffic is the shared weight sweep amortized over the emitted tokens, plus the private context read.

    qKV(b,Lread)=Wbatch(b)bρ+Kread(Lread).q_{\mathrm{KV}}(b, L_{\mathrm{read}}) = \frac{W_{\mathrm{batch}}(b)}{b\,\rho} + K_{\mathrm{read}}(L_{\mathrm{read}}) .

    Then the memory roofline TR/qT \le R/q gives the aggregate ceiling at batch bb,

    T(b,Lread)RqKV(b,Lread),T(b, L_{\mathrm{read}}) \le \frac{R}{q_{\mathrm{KV}}(b, L_{\mathrm{read}})},

    and dividing by the batch gives the per-session rate,

    r(b,Lread)=T(b,Lread)b=ρRWbatch(b)+bρKread(Lread).r(b, L_{\mathrm{read}}) = \frac{T(b, L_{\mathrm{read}})}{b} = \frac{\rho R}{W_{\mathrm{batch}}(b) + b\,\rho\,K_{\mathrm{read}}(L_{\mathrm{read}})} .

    As bb grows, the aggregate TT rises but the per-session rr falls. That tension is the whole serving tradeoff, and it is why a fit-only bound is not enough.

    Usable-batch correction

    The fix is to refuse batches at which a session would crawl. Impose a per-session floor rr_\star, the minimum useful tokens/s/session, and solve r(b)rr(b) \ge r_\star for bb. Replacing the batch-dependent Wbatch(b)W_{\mathrm{batch}}(b) by its shared lower bound WactiveW_{\mathrm{active}} keeps a closed form. Because Wbatch(b)WactiveW_{\mathrm{batch}}(b) \ge W_{\mathrm{active}}, the substitution only weakens the condition, so the implication runs one way, and the closed form is a necessary condition on the admissible batch rather than a sufficient one.

    r(b)rbρR/rWactiveρKread(Lread),r(b) \ge r_\star \quad\Longrightarrow\quad b \le \frac{\rho R / r_\star - W_{\mathrm{active}}}{\rho\,K_{\mathrm{read}}(L_{\mathrm{read}})} ,

    which defines a rate-limited batch

    brate(Lread,r)=ρR/rWactiveρKread(Lread).b_{\mathrm{rate}}(L_{\mathrm{read}}, r_\star) = \left\lfloor \frac{\rho R / r_\star - W_{\mathrm{active}}}{\rho\,K_{\mathrm{read}}(L_{\mathrm{read}})} \right\rfloor .

    The usable batch is whichever gate binds first.

    busable=min ⁣(bmem(Lalloc), brate(Lread,r)).b_{\mathrm{usable}} = \min\!\big(b_{\mathrm{mem}}(L_{\mathrm{alloc}}),\ b_{\mathrm{rate}}(L_{\mathrm{read}}, r_\star)\big).

    Because brateb_{\mathrm{rate}} comes from a necessary condition, busableb_{\mathrm{usable}} is itself an upper bound on the truly admissible batch, and the calculator applies the exact floor test with the true Wbatch(b)W_{\mathrm{batch}}(b) in the next step. This is what stops the “hundred sessions” illusion. As context grows, Kread(L)K_{\mathrm{read}}(L) grows, so brateb_{\mathrm{rate}} falls quickly even while bmemb_{\mathrm{mem}} stays large. The KV slots fit while the useful rate does not.

    Interactive figure: aggregate and per-session ceilings against batch size, with the memory gate and the per-session floor. Enable JavaScript to explore it.

    The bound the calculator uses

    Collecting the pieces, define the usable batch set as the fitting batches that also clear the floor,

    B(Lalloc,Lread,r)={b:1bbmem(Lalloc), 1bRqKV(b,Lread)r},\mathcal{B}(L_{\mathrm{alloc}}, L_{\mathrm{read}}, r_\star) = \left\{ b : 1 \le b \le b_{\mathrm{mem}}(L_{\mathrm{alloc}}),\ \frac{1}{b}\,\frac{R}{q_{\mathrm{KV}}(b, L_{\mathrm{read}})} \ge r_\star \right\},

    and take the best aggregate over that set.

    Tmax(Lalloc,Lread,r)maxbB RWbatch(b)bρ+Kread(Lread).T_{\max}(L_{\mathrm{alloc}}, L_{\mathrm{read}}, r_\star) \le \max_{b \in \mathcal{B}}\ \frac{R}{\dfrac{W_{\mathrm{batch}}(b)}{b\,\rho} + K_{\mathrm{read}}(L_{\mathrm{read}})} .

    This is the KV-aware bound, the main practical formulation. In words, try every batch that fits, reject the ones too slow per session, and for the rest take bandwidth divided by bytes per output token, keeping the best.

    The looser memory-power bound is its corollary, obtained as before by dropping the private term and using the largest fitting batch.

    TmaxρDKalloc(Lalloc)Wactive(1Wresident+OC),D=CR.T_{\max} \le \rho\,\frac{D}{K_{\mathrm{alloc}}(L_{\mathrm{alloc}})\,W_{\mathrm{active}}}\left(1 - \frac{W_{\mathrm{resident}} + O}{C}\right), \qquad D = CR.

    The two stand in a fixed relation, which is the main result of the derivation, written first by name and then in full.

    Tmax    KV-aware bound    memory-power bound    large-memory limitTmax    maxbBRWbatch(b)bρ+Kread(Lread)    ρDKalloc(Lalloc)Wactive(1Wresident+OC)    ρDKalloc(Lalloc)Wactive\begin{aligned} T_{\max} \;&\le\; \text{KV-aware bound} \;\le\; \text{memory-power bound} \;\le\; \text{large-memory limit} \\ T_{\max} \;&\le\; \max_{b \in \mathcal{B}} \frac{R}{\dfrac{W_{\mathrm{batch}}(b)}{b\rho} + K_{\mathrm{read}}(L_{\mathrm{read}})} \\ \;&\le\; \rho\,\frac{D}{K_{\mathrm{alloc}}(L_{\mathrm{alloc}})\,W_{\mathrm{active}}} \left(1 - \frac{W_{\mathrm{resident}} + O}{C}\right) \\ \;&\le\; \rho\,\frac{D}{K_{\mathrm{alloc}}(L_{\mathrm{alloc}})\,W_{\mathrm{active}}} \end{aligned}

    The gap across these terms is the point of the whole derivation. The KV-aware line is the tight, practical bound. The memory-power line shows the memory system’s large theoretical capacity-bandwidth product, and the distance between them comes from private context traffic, expert diversity, and the per-session floor. The final term drops the resident-model factor as well, so the right-hand side is exactly the simplified D/(KW)D/(KW) from the memory-power decode bound, now with K=Kalloc(Lalloc)K = K_{\mathrm{alloc}}(L_{\mathrm{alloc}}) and W=WactiveW = W_{\mathrm{active}}. It is the loosest, most optimistic reading, since the resident model and overhead always claim a real share of CC.

    Single session

    Set b=1b = 1 to recover the latency-style bound for one conversation. There is no batch to amortize the weight sweep over.

    T1,maxRWbatch(1)/ρ+Kread(Lread),T_{1,\max} \le \frac{R}{W_{\mathrm{batch}}(1)/\rho + K_{\mathrm{read}}(L_{\mathrm{read}})},

    and for ordinary decoding (ρ=1\rho = 1) this is just R/(Wbatch(1)+Kread(Lread))R / (W_{\mathrm{batch}}(1) + K_{\mathrm{read}}(L_{\mathrm{read}})). Capacity has dropped out except as the gate that decides whether the model fits at all, consistent with the earlier observation that bandwidth governs single-session speed.

    Speculative decoding

    Speculative decoding lets ρ>1\rho > 1. A draft model proposes several tokens and the target verifies them in one iteration, so ρ\rho is the expected number of accepted tokens per session per verification step, bounded by the draft length γ\gamma as 1ργ+11 \le \rho \le \gamma + 1. The temptation is to multiply throughput by ρ\rho and stop. That is wrong, because the draft model and verification are not free. Their traffic belongs in Wbatch(b)W_{\mathrm{batch}}(b) or in the per-token term. The safe rule is to never scale by ρ\rho without charging the draft cost in the denominator. With both effects included, speculative decoding moves through the same KV-aware formula unchanged.

    The two bounds side by side

    The gap between the memory-power bound and the KV-aware bound is easiest to see as memory traffic. Below, two copies of the same machine decode side by side. Each board is the machine’s usable memory, with the weights packed into an orange container of equal-sized cells and each session’s blue KV cells packed into a small container of its own.

    Each decode iteration must move every byte that its accounting charges, at the same bandwidth on both machines, so the charged cells light up one by one and the board resets when the iteration completes. The left board charges only the shared weight cells, so it resets quickly and its token counter races ahead. The right board also charges the read cells of every session’s context, so its iterations stretch as the batch and the context grow. A real iteration takes milliseconds, so time runs in slow motion here.

    Interactive animation: memory-power accounting and KV-aware accounting decoding side by side on the same machine. Enable JavaScript to watch it.

    Both boards run on the same silicon at the same bandwidth, and only the bookkeeping differs between them. The left counter is the memory-power accounting at the chosen batch, and the right one is what the KV-aware bound admits once private context reads are charged.

    The sliders show the two context lengths at work. The reserved context sets how many cells each session’s container holds and can push the machine past its capacity, so raising it eventually makes the containers stop fitting. The read context sets how many of those cells light up every iteration, and the longer it gets the smaller the weights’ share of each iteration becomes. It follows the reservation at an adjustable fraction, 90 percent by default, with the unread rest capped at 32k, and the sliders for all of this sit under Advanced. Growing the model itself slows both boards down in step, while the orange container eats the room the blue ones need.

    The full memory-power bound goes one step further than the left board. It grows the batch until memory is completely full of KV cells, which is exactly what the default maximize batch mode does, and the line under the left board reports that number. Shrinking the reserved context makes it explode.

    At a 4k reservation about a hundred sessions fit and the bound climbs past 4,000 tok/s on this toy machine, and at a 1k reservation it would pass 17,000. Those numbers are true ceilings and useless forecasts at the same time. What stops a real machine long before then is reading each session’s context, which is exactly the traffic the right board charges.

    Model adapters

    The KV-aware bound is architecture-agnostic, and an architecture enters only through the five quantities. So each model family is captured by a small adapter that supplies them.

    Adapter(M)=[Wresident, Wbatch(b), Kalloc(Lalloc), Kread(Lread), ρ].\mathrm{Adapter}(M) = \big[\, W_{\mathrm{resident}},\ W_{\mathrm{batch}}(b),\ K_{\mathrm{alloc}}(L_{\mathrm{alloc}}),\ K_{\mathrm{read}}(L_{\mathrm{read}}),\ \rho \,\big].

    Three adapters cover the catalog. They handle dense transformers, mixture-of-experts, and hybrid/sliding/recurrent attention.

    Dense transformers

    For a dense model with PtotalP_{\mathrm{total}} parameters at ewe_w bytes each, all weights are touched every step, so the resident footprint and the per-step sweep coincide.

    Wresident=Wbatch(b)=Ptotalew.W_{\mathrm{resident}} = W_{\mathrm{batch}}(b) = P_{\mathrm{total}}\, e_w .

    With NlayersN_{\mathrm{layers}} layers, NkvN_{\mathrm{kv}} key/value heads, head dimension dhd_h, and KV byte widths eKV,storee_{\mathrm{KV,store}} and eKV,reade_{\mathrm{KV,read}}, full-context attention reserves and reads

    Kalloc(L)=2NlayersNkvdheKV,storeL,Kread(L)2NlayersNkvdheKV,readL,K_{\mathrm{alloc}}(L) = 2\, N_{\mathrm{layers}}\, N_{\mathrm{kv}}\, d_h\, e_{\mathrm{KV,store}}\, L, \qquad K_{\mathrm{read}}(L) \approx 2\, N_{\mathrm{layers}}\, N_{\mathrm{kv}}\, d_h\, e_{\mathrm{KV,read}}\, L,

    where the factor 22 counts keys and values. Weight precision and KV precision are independent settings. NVFP4 weights (NVIDIA’s 4-bit floating-point format) do not imply an NVFP4 cache, so ewe_w and eKVe_{\mathrm{KV}} are tracked separately.

    Mixture-of-experts

    An MoE model is where the constant-WbatchW_{\mathrm{batch}} assumption breaks, and fixing it is the single most important adapter correction. Let PtotalP_{\mathrm{total}} be total parameters, PactiveP_{\mathrm{active}} the active parameters per token, EE the number of routed experts, and kk the experts selected per token. Assuming uniformly sized routed experts, each routed expert holds

    pexpert=PtotalPactiveEk,p_{\mathrm{expert}} = \frac{P_{\mathrm{total}} - P_{\mathrm{active}}}{E - k},

    and the always-on remainder (dense trunk, shared experts, embeddings, attention) is

    Pfixed=Pactivekpexpert.P_{\mathrm{fixed}} = P_{\mathrm{active}} - k\, p_{\mathrm{expert}} .

    The naive model assumes a batch touches the same active experts every session, keeping WbatchW_{\mathrm{batch}} constant. That is false. Independent sessions route to different experts, so a larger batch touches more distinct experts. With bρb\rho token-routings, each independently missing a given expert with probability 1k/E1 - k/E, the expected number of distinct experts touched is

    m(bρ)=E(1(1kE)bρ),m(b\rho) = E\left(1 - \left(1 - \frac{k}{E}\right)^{b\rho}\right),

    and the per-iteration shared traffic is the fixed part plus the touched experts.

    Wbatch(b)=ew[Pfixed+pexpertm(bρ)].W_{\mathrm{batch}}(b) = e_w\big[\, P_{\mathrm{fixed}} + p_{\mathrm{expert}}\, m(b\rho) \,\big].

    At bρ=1b\rho = 1 this reduces to the active-parameter footprint, and as bρb\rho \to \infty it saturates at all EE experts. This rising Wbatch(b)W_{\mathrm{batch}}(b) is why MoE batching does not amortize for free, and why the MoE rows in the worked table reach their throughput optimum at modest batch sizes.

    One caveat applies here. Every other traffic term in the bound is a deliberate under-estimate of real traffic, which is what makes R/qR/q a true ceiling. The expert count m(bρ)m(b\rho) is different. It is an expectation under independent, uniform routing rather than a lower bound. Real routing is correlated, since load-balancing losses push toward uniform while hot experts and topically similar sessions pull the other way, and correlated routing touches fewer distinct experts than the formula predicts. In that case the modeled traffic overstates the actual traffic, and the computed ceiling can sit below the true one. When a guaranteed ceiling is required, replace m(bρ)m(b\rho) by its minimum kk, which replaces Wbatch(b)W_{\mathrm{batch}}(b) by WactiveW_{\mathrm{active}}. The expectation form is the better estimate, the floor form is the safe bound.

    Hybrid, sliding, and recurrent attention

    Models with local or sliding-window attention, compressed or latent attention, or linear/recurrent state must not use the full-KV formula blindly, because their cache does not grow linearly in LL everywhere. Split both KV terms into global, local, and fixed-state parts.

    K(L)=Kglobal,(L)+Klocal,(L)+Kstate,,{alloc,read}.K_{\bullet}(L) = K_{\mathrm{global},\bullet}(L) + K_{\mathrm{local},\bullet}(L) + K_{\mathrm{state},\bullet}, \qquad \bullet \in \{\mathrm{alloc}, \mathrm{read}\}.

    A simple read approximation with sliding-window width ww is

    Kread(L)=κglobalL+κlocalmin(L,w)+Kstate.K_{\mathrm{read}}(L) = \kappa_{\mathrm{global}}\, L + \kappa_{\mathrm{local}}\,\min(L, w) + K_{\mathrm{state}} .

    Full-attention layers pay for the whole context, sliding-window layers pay only up to the window, and recurrent or latent state adds a fixed or slowly growing term. The same shape covers Gemma-style local/global attention and DeepSeek-style compressed/sparse attention, with only the coefficients changing.

    Calculator procedure

    The bound is now ready to compute. Because the same computation runs for every hardware-and-model pair, it is worth stating once as a procedure.

    The inputs are a hardware row (C,R,O)(C, R, O), a model adapter (Wresident,Wbatch(),Kalloc(),Kread(),ρ)(W_{\mathrm{resident}}, W_{\mathrm{batch}}(\cdot), K_{\mathrm{alloc}}(\cdot), K_{\mathrm{read}}(\cdot), \rho), and workload assumptions (Lalloc,Lread,r)(L_{\mathrm{alloc}}, L_{\mathrm{read}}, r_\star).

    1. Compute the resident margin CWresidentOC - W_{\mathrm{resident}} - O. If it is negative, the model does not fit, so stop.
    2. Compute Kalloc(Lalloc)K_{\mathrm{alloc}}(L_{\mathrm{alloc}}) and the memory-fit batch bmemb_{\mathrm{mem}}.
    3. For each integer batch 1bbmem1 \le b \le b_{\mathrm{mem}}, compute Wbatch(b)W_{\mathrm{batch}}(b) and qKV(b,Lread)q_{\mathrm{KV}}(b, L_{\mathrm{read}}).
    4. Compute the aggregate ceiling R/qKVR / q_{\mathrm{KV}} and the per-session ceiling R/(b,qKV)R / (b, q_{\mathrm{KV}}) at each batch.
    5. Keep the batches whose per-session ceiling is at least rr_\star.
    6. Among the kept batches, choose the one with the largest aggregate ceiling, and report it as the KV-aware result with its batch as bb^\star.
    7. Separately compute the memory-power ceiling for orientation.

    The output is a stack of gates, and the right phrasing depends on which gate bound.

    StateMeaning
    Resident fitThe model plus overhead fits in memory
    Session fitAt least one reserved-context session fits
    Floor fitSome fitting batch clears rr_\star
    No floorSessions fit, but no batch clears rr_\star

    The common invalid reading is that fitting in memory implies serving usefully. A model can pass resident fit and session fit and still have an empty usable batch set, because every fitting batch is below the floor. The honest report for that case is “fits, but no batch satisfies the floor”, a distinct verdict from a true fit failure. Keeping the two apart is the reason the floor gate exists.

    Worked examples

    Now check the theory against real hardware. Consider two 128 GB machines, one bandwidth-rich and one bandwidth-poor, which isolate the effect of RR at fixed CC.

    Two machines

    NVIDIA’s DGX Spark, a small desktop AI machine, carries 128 GB of LPDDR5x unified memory at 273 GB/s. An Apple M5 Max with a 40-core GPU reaches 614 GB/s and is configurable to 128 GB of unified memory. At equal capacity their memory powers are

    HardwareCCRRD=CRD = CR
    DGX Spark128 GB273 GB/s34,944 GB2/s\mathrm{GB}^2/\mathrm{s}
    Apple M5 Max 128GB128 GB614 GB/s78,592 GB2/s\mathrm{GB}^2/\mathrm{s}

    Both bandwidth numbers are catalog figures rather than measured sustained rates, so the earlier caveat applies. If one machine sustains a larger share of its spec than the other, the comparison will make the other machine look better than it really is.

    Three models

    Three MoE models, modeled from their published cards as adapter parameters, without re-measurement.

    • Qwen3.6-35B-A3B, with 35B total / 3B active parameters, E=256E = 256 routed experts, k=8k = 8 routed (plus one shared) per token, and weights quantized NVFP4.
    • Gemma 4 26B-A4B-it, with 26B total / 4B active, E=128E = 128, top-8 routing, hybrid local/global attention, and NVFP4 weights.
    • DeepSeek V4 Flash (DS4), with 284B total / 13B active, E=256E = 256 routed plus one shared, k=6k = 6 per token, million-token context via compressed/sparse attention, and weights at a Q2-style mixed quantization.

    All rows use the Local Frontier defaults, namely reserved context Lalloc=100,000L_{\mathrm{alloc}} = 100{,}000, active context Lread=32,000L_{\mathrm{read}} = 32{,}000, per-session floor r=20r_\star = 20 tok/s/session, ordinary decoding ρ=1\rho = 1, runtime overhead O=8O = 8 GB, and the memory roofline only. The numbers are memory-side upper bounds from the simplified adapters, to be read as ceilings for comparing hardware.

    Results

    HardwareModelSingle-sessionbb^\starKV-aware aggregateMemory-power ceiling
    DGX SparkQwen3.6-35B-A3B\le 149 tok/s17\le 345 tok/s\le 18.2k tok/s
    DGX SparkGemma 4 26B-A4B-it\le 120 tok/s16\le 333 tok/s\le 23.7k tok/s
    DGX SparkDeepSeek V4 Flash (Q2)\le 83 tok/s7\le 154 tok/s\le 13.1k tok/s
    Apple M5 Max 128GBQwen3.6-35B-A3B\le 336 tok/s50\le 1,006 tok/s\le 41.0k tok/s
    Apple M5 Max 128GBGemma 4 26B-A4B-it\le 271 tok/s66\le 1,326 tok/s\le 53.3k tok/s
    Apple M5 Max 128GBDeepSeek V4 Flash (Q2)\le 188 tok/s22\le 446 tok/s\le 29.5k tok/s

    Two things stand out. First, with capacity held equal, the higher M5 Max bandwidth lifts the single-session ceilings in proportion to RR, and the batched ceilings by even more, because the extra bandwidth also lets more sessions clear the per-session floor. The bandwidth-rich machine wins exactly where the theory says it should, in batched throughput. Second, the memory-power column sits one to two orders of magnitude above the KV-aware column. That gap is the cost of private context traffic and expert diversity, and showing it is the point of the derivation.

    Forced concurrency

    What if concurrency is fixed by policy rather than chosen at the floor-satisfying optimum? On DGX Spark, pushing past bb^\star buys aggregate throughput at the cost of per-session rate.

    ModelBatch bbAggregate ceilingPer-session ceiling
    Qwen3.6-35B-A3B32\le 394 tok/s\le 12.3 tok/s/session
    Qwen3.6-35B-A3B64\le 478 tok/s\le 7.5 tok/s/session
    Gemma 4 26B-A4B-it32\le 430 tok/s\le 13.4 tok/s/session
    Gemma 4 26B-A4B-it64\le 575 tok/s\le 9.0 tok/s/session

    This is the serving tradeoff in numbers. For DGX Spark under these assumptions, 32 and 64 sessions are too high if the goal is around 20 tok/s/session, exactly the regime the usable-batch correction is built to reject, and the reason bb^\star for these models settles near 16.

    A sanity check against a real run

    A reported DGX Spark run served Gemma at concurrency 16 at roughly 16 to 18 tok/s/session, an aggregate of 16×16=25616 \times 16 = 256 to 16×18=28816 \times 18 = 288 tok/s. The KV-aware aggregate ceiling for Gemma at this batch is 333\le 333 tok/s, so observed throughput is

    25633377%to28833386%\frac{256}{333} \approx 77\% \quad\text{to}\quad \frac{288}{333} \approx 86\%

    of the simplified ceiling. That is close enough to suggest the implementation is near the memory-side roofline. It does not prove the quantization is optimal. The bound omits compute, scheduler behavior, kernel details, and exact cache traffic, and proving optimality would require profiler evidence of bandwidth saturation with no compute, scheduler, or CPU stalls. A bound this close simply means there is little memory headroom left to capture.

    Omitted rooflines

    The memory roofline is one limit among several. Real throughput is the minimum over all of them,

    Tmaxmin(Tmemory, Tcompute, Tkernel, Tscheduler, Tinterconnect),T_{\max} \le \min\big(T_{\mathrm{memory}},\ T_{\mathrm{compute}},\ T_{\mathrm{kernel}},\ T_{\mathrm{scheduler}},\ T_{\mathrm{interconnect}}\big),

    and the memory term we computed can be undercut by compute throughput and tensor-core utilization, dequantization kernels, attention kernels and KV layout, prefill/decode phase mixing, scheduler overhead and request churn, CPU and PCIe involvement, multi-GPU communication, allocator fragmentation, thermal and power limits, tokenization and sampling, speculative rejection rates, and prefix-cache hit rates. Each belongs as its own limit term. The memory model remains useful because it makes the first unavoidable ceiling explicit and cheap to compute, and because for memory-bound decode it is usually the binding one.

    One recurring caution applies to models with recurrent or linear state. A model with tiny fixed state and tiny private read traffic produces an enormous memory-side aggregate at high concurrency, because almost nothing in the denominator grows with the batch. That is precisely the signal that compute, kernel, scheduler, and recurrent-state details must be added before the aggregate number is treated as realistic. The memory bound describes the best case the hardware allows, and reaching it is the implementation’s job.

    Cheat sheet

    This section collects the whole formulation in one place, so it can be read on its own. A machine is three numbers and a model adapter is five, with the workload adding three assumptions.

    SymbolMeaning
    CC, RR, OOUsable memory capacity, sustained memory bandwidth, runtime overhead
    WresidentW_{\mathrm{resident}}Full resident weight footprint, which must fit in memory
    Wbatch(b)W_{\mathrm{batch}}(b)Shared weight traffic per iteration, equal to WresidentW_{\mathrm{resident}} for dense models and growing with bb for MoE
    Kalloc(Lalloc)K_{\mathrm{alloc}}(L_{\mathrm{alloc}})KV memory reserved per session
    Kread(Lread)K_{\mathrm{read}}(L_{\mathrm{read}})Private context traffic per output token
    ρ\rhoTokens emitted per session per iteration, one for ordinary decoding
    LallocL_{\mathrm{alloc}}, LreadL_{\mathrm{read}}Reserved and average read context, LreadLallocL_{\mathrm{read}} \le L_{\mathrm{alloc}}
    rr_\starMinimum useful tokens/s per session

    Everything descends from the memory roofline. If each output token must move at least qq bytes and the machine delivers at most RR bytes per second, then

    TmaxRq.T_{\max} \le \frac{R}{q} .

    The first gate is whether sessions fit. Load the weights, reserve the overhead, and divide what is left by the per-session KV allocation.

    bmem(Lalloc)=CWresidentOKalloc(Lalloc).b_{\mathrm{mem}}(L_{\mathrm{alloc}}) = \left\lfloor \frac{C - W_{\mathrm{resident}} - O}{K_{\mathrm{alloc}}(L_{\mathrm{alloc}})} \right\rfloor .

    Per-token traffic is the shared weight sweep amortized over the batch plus the private context read, which no batch size amortizes.

    q(b,Lread)=Wbatch(b)bρ+Kread(Lread).q(b, L_{\mathrm{read}}) = \frac{W_{\mathrm{batch}}(b)}{b\,\rho} + K_{\mathrm{read}}(L_{\mathrm{read}}) .

    The roofline gives the aggregate ceiling at batch bb, and dividing by the batch gives the per-session rate. The aggregate rises with bb while the per-session rate falls.

    T(b)Rq(b,Lread),r(b)=T(b)b=ρRWbatch(b)+bρKread(Lread).T(b) \le \frac{R}{q(b, L_{\mathrm{read}})}, \qquad r(b) = \frac{T(b)}{b} = \frac{\rho R}{W_{\mathrm{batch}}(b) + b\,\rho\,K_{\mathrm{read}}(L_{\mathrm{read}})} .

    Imposing the floor rr_\star rejects the batches where every session crawls, and the usable batch is whichever gate binds first.

    brate(Lread,r)=ρR/rWactiveρKread(Lread),busable=min(bmem, brate).b_{\mathrm{rate}}(L_{\mathrm{read}}, r_\star) = \left\lfloor \frac{\rho R / r_\star - W_{\mathrm{active}}}{\rho\,K_{\mathrm{read}}(L_{\mathrm{read}})} \right\rfloor, \qquad b_{\mathrm{usable}} = \min\big(b_{\mathrm{mem}},\ b_{\mathrm{rate}}\big) .

    The usable batch set holds the batches that fit and clear the floor, and the KV-aware bound is the best aggregate over it.

    B={b:1bbmem, Rbq(b,Lread)r},TmaxmaxbBRq(b,Lread).\mathcal{B} = \left\{ b : 1 \le b \le b_{\mathrm{mem}},\ \frac{R}{b\, q(b, L_{\mathrm{read}})} \ge r_\star \right\}, \qquad T_{\max} \le \max_{b \in \mathcal{B}} \frac{R}{q(b, L_{\mathrm{read}})} .

    Setting b=1b = 1 in the same formula gives the single-session ceiling. Dropping the private context term and taking the largest fitting batch gives the looser memory-power ceiling,

    TmaxρDKalloc(Lalloc)Wactive(1Wresident+OC),D=CR,T_{\max} \le \rho\,\frac{D}{K_{\mathrm{alloc}}(L_{\mathrm{alloc}})\,W_{\mathrm{active}}}\left(1 - \frac{W_{\mathrm{resident}} + O}{C}\right), \qquad D = CR ,

    and the three levels always order the same way.

    actual throughput    KV-aware bound    memory-power bound.\text{actual throughput} \;\le\; \text{KV-aware bound} \;\le\; \text{memory-power bound} .

    Every number these formulas produce is an upper bound built from memory capacity and bandwidth alone. Real implementations land below it, and compute, software overhead, and interconnects can only lower the ceiling further.

    When reporting a result, always state the assumptions that move it, namely LallocL_{\mathrm{alloc}}, LreadL_{\mathrm{read}}, ρ\rho, rr_\star, the weight precision, and the KV-cache precision or attention adapter. Without them, a single tok/s number is not reproducible.

    To see these bounds computed live for hundreds of audited model profiles against a catalog of local hardware, try the Local Frontier calculator.