---
title: "Local AI will favor LPDDR and MoE"
date: 2026-08-10
canonical: https://solmaz.io/x/2086715364642365804/
x_url: https://x.com/onusoz/status/2086715364642365804
license: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
---

This.

LPDDR chips are cheaper to produce and run

GDDR/HBM will likely keep being more expensive

Most consumer GPUs will converge on a GB10 like form factor

As much as us hobbyists love to project this ideal of running a GPU cluster at home, most working people will prefer smaller form factors, and will not want to pay hundreds of $$$ in electricity bills every month

DGX Spark/GB10 runs at around 90-150 Watts
RTX Pro 6000 runs at 600 Watts FOR THE GPU ALONE, and can cost 3-5x more than GB10. Despite having 25% less memory capacity than GB10...

Looking at this, LPDDR will be orders of magnitude more commonplace at home

Architectures  will develop accordingly. Future local AI will be dominated by MoE and similar architectures which leverage mid-sized models with smaller number of active parameters

That is why Qwen3.x-35B-A3B is a more useful model on the Spark than Qwen3.x-27B, despite the latter being a better model. Same for Gemma

I can run A3B at 60 decode tok/s single session or 6x20 decode tok/s in parallel, whereas 27b only reaches 1/3rd of that

Future of local AI is DDR/LPDDR and MoE/adjacent architectures, for the average person

*Quotes a post by @TheAhmadOsman (https://x.com/TheAhmadOsman/status/2086195772890960159); its text is omitted here because it is not covered by this site's license.*
