AI model reliability
Mentioned in
- We have AGI, but it cannot write an abstract
- Measuring model harness sensitivity
- Standardized test harnesses for model benchmarks
- Use workflows when models cannot follow rules
- Fable made Alman possible
- Anthropic is optimizing for general knowledge work
- Inference scaling can reduce coding model quality
- Stop defaulting to weak coding agents for serious work
- GPT-5.2 xhigh feels like a careful systems debugger
- Predictions by Anthropic Researchers