Post @onusoz · /2026/09/04 · 06:14 AM View on not only this, but evaluating model output will be a chore for most knowledge workers every company will have to benchmark their agents on their use cases for their customers, across all sectors @cancelik · Sep 3, 2026 generic benchmarks will die and personal benchmarks will rise, or at least should. gemini beats astra on deepswe, astra gets same score with sol on artificial analysis and astra gets 99% on arc-agi. benchmarks are pure noise at this point. Show more
@onusoz · /2026/09/04 · 06:14 AM View on not only this, but evaluating model output will be a chore for most knowledge workers every company will have to benchmark their agents on their use cases for their customers, across all sectors @cancelik · Sep 3, 2026 generic benchmarks will die and personal benchmarks will rise, or at least should. gemini beats astra on deepswe, astra gets same score with sol on artificial analysis and astra gets 99% on arc-agi. benchmarks are pure noise at this point. Show more