Post
-
I have made an update to my theoretical upper bound calculation to also predict prefill speed Prefill relaxes the assumption we make for decode, that it is only be memory bottlenecked. So prefill can be both compute or memory bottlenecked. I use the FLOP limits reported by hardware producers for the estimates: These estimates will also be available in ourmodels.cc for indexed model and hardware in a couple days, once a long running job finishes Blog post: solmaz.io/llm-performance-upper-bounds