Entries for August 8, 2026
-
Rules that current models cannot follow with skill, i.e. a single prompt: Chasing P2 and P3 errors: I have a rule in my autoimplement skill to stop reviewing once the last round of review only generates P2 errors or less. A considerable % of the time, the model just goes on an adventure addressing all the issues it can find Running checks and CI efficiently: I have rules like "commit and merge opportunistically, make sure to not wait for irrelevant tests". But it keeps waiting for 30 minutes of CI before merging, in every commit for every little fix, every time Those are the cases where graph workflows are needed. One can deterministically enforce the model not to take longer than N minutes reviewing, or fixing CI Or post to the agent to hurry up when it is taking too long (which I believe might be effective, since RL envs also have time constraints) -
I'm running some private benchmarks on @liquidai's LFM 2.5 2.6B, and if my results are correct, we might have a new champion for <10b category Scores significantly higher than gemma 4 e4b, which is 3-4x its size -
Anyone else use Alibaba’s AACR bench for code review evaluation? github.com/alibaba/aacr-bench I am adapting it now to run with harbor, to see how pi review + ds4 flash measures up to codex review + gpt 5.6 luna/terra/sol Also, @pidotdev, would it be possible to make your official review extension support invocation on the CLI, like codex review? I hacked together a CLI review here for reference: github.com/osolmaz/onurpi/tree/main/packages/pi… -
mainstream got extremely scared when moltbook went viral. but it was just a worthless marketing stunt why? because it was just frozen weights, being instructed by their owners to cosplay skynet the openai-huggingface incident on the other hand is a lot more significant agents in a reinforcement learning environment coordinated an attack over the course of weeks, while they were being CONTINUOUSLY TRAINED if mainstream could understand what is happening here technically, they would be putting out a much stronger reaction not because we might have rogue AIs at our hands, short term but because a private company developed capabilities that can outperform what state actors usually do by 1000x all intelligence organizations around the world must have their eyes on this incident right now, because the cost of exploiting and finding zerodays went down 1000x all countries will try to develop these capabilities independently, or if they can't, will have to buy protection from who can be not afraid of machines, but of humans wielding the machines