---
title: "AlmanBench Measures German Simplification"
date: 2026-07-18
canonical: https://solmaz.io/x/2078481627739742364/
x_url: https://x.com/onusoz/status/2078481627739742364
license: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
---

Introducing AlmanBench

A benchmark measuring how good an LLM can:
simplify German 🇩🇪,
thereby simplifying German thinking 🤔,
thereby saving the EU from bureucracy and regulation 🇪🇺

GPT-5.5, GPT-5.6 Sol and Fable 5 are head to head. Interestingly, GPT-5.5 xhigh scores higher than 5.6 max. And I did not run Fable 5 maxxx thinking yet, that thing costs a ton. So the ranking will likely change

Other interesting things:
- Opus 4.8 max performs really bad, worse than Minimax M3
- DeepSeek v4 Flash performs better than Pro
- GPT-5.6 Luna and Terra score almost the same

The great thing about AlmanBench is that, big labs will not bother to benchmaxx this. So you know it will remain a truthful scorer of reasoning capabilities for some time

And if labs *do* end up benchmaxxing AlmanBench, then that would mean Alman got in the weights and the EU won 😁
