---
title: "Standardized Test Harnesses for Model Benchmarks"
date: 2026-08-23
canonical: https://solmaz.io/x/2091434969151267162/
x_url: https://x.com/onusoz/status/2091434969151267162
license: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
---

We need to normalize measuring and judging models against a standardized test harness

"Oh but model X performs best in their own proprietary harness"

I could not care less. When I take exams, I go to the standardized classroom, get the standardized pencil and exam sheet, and have to solve it under 2 hours

This system arose because we have a LOT of people to test

Guess what? We now have a LOT of models, and they are multiplying by the day

"Oh but model X performs substantially better in ARC-AGI-3 with a custom harness"

I don't care... Then imbue model X with enough knowledge so that it can reconstruct that harness on the spot

The main harness could be mini-swe-agent, terminus 2, vanilla pi or something along those lines

It needs to be simple, and stay roughly the same over time

There is already too much complexity in the benchmarking space right now, and I feel like not enough people are putting their feet down to cut away some part of it
