Post
-
LFM 2.5-2.6B by @liquidai just launched and it punches above its weight! It can run 32 sessions (and more) in parallel with hundreds of output tokens per second aggregate throughput, on the DGX Spark! And this is just the base vLLM config on release date, I expect it to be optimized a lot more!Image hidden