Entries for August 27, 2026
-
glm 5.3 flash at an antirez-style asymmetric q2 quant could plausibly fit on a single dgx spark at 100~110 GB but with 18b active params vs ds4 flash’s 13b, theoretical decode throughput is only ~70% of that of ds4 flash on the same hardware that is 28 tok/s ds4 vs 19~20 tok/s glm, without speculative decoding, at 32k token filled context afaik it hasn't been quantized in that style yet, but if it were, we would expect such throughput ratio between those models