Entries for September 11, 2026
-
4 instances of DeepSeek V4.1 Flash played Settlers of Catan against each other The whole run took 4.5 hours, 77 turns, 634 API calls and cost 5.7 usd (much slower than an average human game, despite 200 tok/s on novita) Red took the lead on turn 13 and was breaking out. But others formed an alliance against red and held it off 28 turns till the end Now running 2 ds41 flash against 2 luna to see which one will win I've built a Catan simulator and a pi-based harness for the models to be able to interact with the game Would you like to see me RL a small model to play and beat against SOTA flagship models? If so, reply below 👇 You can see the full game state and session logs on Hugging Face as well Repo: github.com/osolmaz/catanarchy Bucket with game data: -
-
-
In the Jacobian conjecture and other recent cases (if not all), agents acted like counterexample monkeys brute-forcing their way into a disproof in a way that does not show the same level of eloquence as a human, by human standards The bar has shifted. We are not impressed anymore by 100 year old problems being solved through millions of $$$ in compute, mathematical equivalent of throwing dynamite at a problem until it breaks My bet for OpenAI Hodge result is yet another counterexample disproof (but apparently Hodge is harder to disprove by brute force because finding a candidate counterexample isn’t enough. you also have to prove no algebraic cycle could ever generate it. I haven't studied this problem before, so take it with a grain of salt) This means we might have a new way to "prove" theorems: If you spend $10m on an agent swarm and they can't disprove it, there is a high chance it might be true :P In 1 year from now, we will have disproven all the low-hanging conjectures, and the remaining set will likely have a higher share of true conjectures than false ones :P -
Hold that thought, there is still hope: