---
title: "glm 5.3 flash at an antirez-style asymmetric q2 quant could plausibly fit on a single dgx spark..."
date: 2026-08-27
canonical: https://solmaz.io/x/2092844182650114452/
x_url: https://x.com/onusoz/status/2092844182650114452
license: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
---

glm 5.3 flash at an antirez-style asymmetric q2 quant could plausibly fit on a single dgx spark at 100~110 GB

but with 18b active params vs ds4 flash’s 13b, theoretical decode throughput is only ~70% of that of ds4 flash on the same hardware

that is 28 tok/s ds4 vs 19~20 tok/s glm, without speculative decoding, at 32k token filled context

afaik it hasn't been quantized in that style yet, but if it were, we would expect such throughput ratio between those models

*Quotes a post by @Zai_org (https://x.com/Zai_org/status/2092616204787626030); its text is omitted here because it is not covered by this site's license.*
