Post
-
DeepSeek V4.1 Flash may be benchmark-contaminated
People who have been running DeepSeek V4.1 Flash on private benchmarks: how does it compare against V4 Flash 0731? To V4 Pro? I have a few data points now and it seems like the model is a little bit benchmaxxed or contaminated (specifically on terminal bench 2.1), and its capabilities might not be so close to GPT 5.6 Sol like the official report implies This is a new architecture that has been trained from scratch, and they could only have released it once it surpassed their previous models enough in capability They seem to have slightly changed strategy, to release as early as possible without degrading quality, due to the increased rate of competition on all fronts And looking at V4, we should expect weight updates that will carry the performance of this architecture even further Looking at their release cadence, maybe around oct 16? on a friday?