Post
  1. Portrait of Onur Solmaz

    AlmanBench Measures German Simplification

    @onusoz · /2026/07/18 · View on
    Introducing AlmanBench A benchmark measuring how good an LLM can: simplify German 🇩🇪, thereby simplifying German thinking 🤔, thereby saving the EU from bureucracy and regulation 🇪🇺 GPT-5.5, GPT-5.6 Sol and Fable 5 are head to head. Interestingly, GPT-5.5 xhigh scores higher than 5.6 max. And I did not run Fable 5 maxxx thinking yet, that thing costs a ton. So the ranking will likely change Other interesting things: - Opus 4.8 max performs really bad, worse than Minimax M3 - DeepSeek v4 Flash performs better than Pro - GPT-5.6 Luna and Terra score almost the same The great thing about AlmanBench is that, big labs will not bother to benchmaxx this. So you know it will remain a truthful scorer of reasoning capabilities for some time And if labs *do* end up benchmaxxing AlmanBench, then that would mean Alman got in the weights and the EU won 😁