Entries for September 12, 2026
-
A threshold was crossed with DeepSeek V4.1 Flash, similar to when Claude Code launched, or the last Christmas of Agents. It's so fast and cheap, 200-250 tok/s on Novita Coupling this model with one of the bigger flagship models can get one much faster to the finishing line in any work So bullish for everyone to follow similar architectures -
People who have been running DeepSeek V4.1 Flash on private benchmarks: how does it compare against V4 Flash 0731? To V4 Pro? I have a few data points now and it seems like the model is a little bit benchmaxxed or contaminated (specifically on terminal bench 2.1), and its capabilities might not be so close to GPT 5.6 Sol like the official report implies This is a new architecture that has been trained from scratch, and they could only have released it once it surpassed their previous models enough in capability They seem to have slightly changed strategy, to release as early as possible without degrading quality, due to the increased rate of competition on all fronts And looking at V4, we should expect weight updates that will carry the performance of this architecture even further Looking at their release cadence, maybe around oct 16? on a friday? -
Another deepseek v4.1 flash issue, model started to refuse work around 700k token line, saying "context is gone", and it has no context repeatedly deepseek advertises 1m context, but the usable context with this model is probably still around 200-300k like other models. I've set that as max before compactionImage hidden