Entries for 2026
-
-
-
When I started building this, the docker space was actually free as well. But due to some bad actors exploiting the free tier at a mindblowing scale, free docker space had to be retracted from the free tier ☹️ However, you can still run ML Claw locally, but with your state backed up to a Hugging Face bucket for free, up to 100 GB of free private storage and 8 TB of public storage! Repo if you would like your agent to scan it before you run the command: github.com/osolmaz/mlclaw -
-
I asked ML Claw to name itself. Of course in German/Alman Lame suggestions: Ulf, Strich, Fritz, Lex I intervened and called it GoePT Goethe rolling in his graveImage hidden -
Fable just seems to know what you want, and incredibly empathetic model AGI is definitely here This model can do anything the average human does and more, given the right context I don't like Anthropic's marketing team, but you've got to hand it to them. They reached there before OpenAI did, despite not having a head-start Now, I can't wait to be able to run a model of this caliber locally 🚀 -
Fable just seems to know what you want, and incredibly empathetic model AGI is definitely here This model can do anything the average human does and more, given the right context I don't like Anthropic's marketing team, but you've got to hand it to them. They reached there before OpenAI did, despite not having a head-start Now, I can't wait to be able to run a model of this caliber locally 🚀 -
-
-
-
I have been working on this for 4 years Every time, I shelved it, because the models were not good enough I tried it with Claude 2 I tried with Gemini 2.5 Pro I tried it with Claude Opus 4 None of them were good enough Until Fable came along Fable achieved perfect score in my hand curated benchmark for Alman. It also one-shotted the landing page, features and translations you see on alman.ai. It is busy creating the golden dataset right now for training I am also happy to announce AlmanBench, a benchmark that measures how well models can translate from German to Alman. I will be posting about performance when a new model drops This idea should have died with me. But thanks to AI, I can unleash my madness onto the world 😈 -
The EU 🇪🇺 is broken. It is drowning in regulation The reason for that is simple Germany 🇩🇪 is the dominant economy of the EU And German 🇩🇪 is the most rule-ridden language in the world Coincidence? I think not. The language people speak determines their thinking and their fate. The Sapir-Whorf hypothesis holds The solution is simple: Fix German, Save the EU 💪 I will SAVE EUROPE by simplifying German. With the help of AI. Once and for all 🫡 I will use OpenClaw 🦞 running on Hugging Face 🤗 to train a model that simplifies German The agent harness is called ML Claw. It has full access to GPUs, jobs, sandboxes and datasets on Hugging Face: github.com/osolmaz/mlclaw My new German dialect is called Alman: alman.ai I will be live-tweeting my Quixotic adventure as it unfolds, in this thread My goal is to show that you can run autoresearch loops on Hugging Face infra on free private Spaces, while getting access to SOTA open models like GLM 5.2. And even use your own Codex subscription! And it does not have to be ML. You can just use free Hugging Face Spaces for your OpenClaw agents. It's free compute! Bookmark to be able to find this thread later on 👇Image hidden -
My friend Poli at @huggingface is the literal 1st to find out about new cool model drops in AI (he has his setup) He is the definition of *alpha* in AI When there is a new model drop, his agents autonomously create a demo around it, for better visibility and interactivity (with human curation of course) And now he created the @HuggingApps account to share these demos live, here on Twitter So if you follow this account, you will be getting the state-of-the-art directly from the publisher. Not weeks, but hours after! Give @HuggingApps and @multimodalart a follow! -
@TheAhmadOsman said it best opensourceaimustwin.com -
A lot of criticism coming towards local models, and the money people are spending to run them Some of these criticisms are valid. No, the layperson will not buy a DGX station, nor spend $50k on a rig They will spend max $3k on a computer, and if their work really necessitates it, up to $10k But in some of these criticisms, I see a lack of first-principles thinking Local models might be shittier now compared to the ones you buy from the cloud The question is, do you WANT to LIVE in a world where you don't have a choice but to RENT your intelligence? From a duopoly that can extort you for the last cent in your wallet, in the long run? I personally do not like that idea. So open source AI HAS to win and keep on winning, continuously and permanently. There is no other choice There is also a preferential belief that AGI/ASI will be reached, but then we will NOT be able to use that to make local models work efficiently??? Like, can you believe the rate of optimization and compression that has been happening to models in the last 3 months? Would you have believed that a o4-mini level model could run in your phone and o4 in your desktop, if we had told you 1 year ago When we are finally there, we will probably look back at this moment and think, "wow, we had not even started to see the real gains from optimization" Local and open source AI HAS to win, and it is up to US So STOP complaining and get to work -
Fable 5 replaces GPT 5 Pro as the Oracle. And it is the go-to model for writing docs So I now have a $ use-fable skill to use whenever I write a README for example, inside codex, through acpx. Saves me from having to switch to another CLI github.com/osolmaz/tools/…Image hidden -
How to get scraped
Note: this post is AI-assisted. It was written with Claude Fable 5 through Cursor, in the same working session that implemented everything it describes.
The internet is busy building walls against AI crawlers. Cloudflare blocks them by default now, publishers are suing, and every other blog post on the topic explains how to keep the crawlers out. This post is the opposite. I spent a day optimizing this site so that AI labs can scrape it as easily as possible, and this is the complete instruction set, so you can do the same to your own site in an afternoon.
My reasoning is that I will die and the weights might not.
Agent-famous
Every frontier model has read a compressed version of the public internet. When your writing is in the training corpus, models absorb your ideas and a faint imprint of how you think. When it is excluded, you don’t exist to them, and “them” increasingly means the layer through which other people experience the internet.
I think being agent-famous is becoming as important for a career as being human-famous. Agent-famous means the models know who you are without having to search. Ask a model about finite element exterior calculus or about keeping AI agents from littering your repo with Markdown files, and if it volunteers your name unprompted, you are in the canon. Search results get you cited when someone happens to look; the weights get you consulted by default. People already pick libraries, tools, and even consultants by asking an agent first, so the difference is starting to pay rent.
Whether your site actually makes it into a training set is decided by the labs, and no amount of optimization changes that. What you control is whether there is any excuse to skip you. A crawler that hits a JavaScript wall, an ambiguous license, or a Cloudflare challenge page will move on, and nobody at the lab will ever know what was behind it. The whole game is removing those excuses.
The view without JavaScript
Start by looking at your site the way a crawler does, which means without JavaScript. I did this last week during a platform migration and found something embarrassing. Every equation on this blog going back to 2017 was rendered client-side by MathJax, so a human with a browser saw beautiful math while a crawler saw raw TeX soup, or nothing. My most substantial technical writing had been invisible to every scraper for eight years.
The fix was moving math rendering to build time (KaTeX in my case), so equations ship as static HTML. The general rule covers more than math. Anything that only exists after JavaScript runs, whether charts, tabs, or content behind a “read more”, does not exist for most of the pipelines that assemble training data. Static HTML is the baseline.
The one-command download
Extractors like trafilatura are decent at pulling article text out of HTML, but why make the labs work for it? Serve your content as plain text yourself:
- Every page on this site now has a raw markdown mirror. Append
.mdto any URL, like /about.md. - /llms.txt is a markdown index of all 773 documents with titles and dates, following the llmstxt.org convention.
- /llms-full.txt is the entire site in one plain-text file, about a megabyte. Each document carries a small frontmatter block with its title, date, canonical URL, and license, and documents are separated by a delimiter line.
So the whole blog is one command:
curl https://solmaz.io/llms-full.txtIf your site is built with a static site generator, all of this is a few small templates over content you already have. Mine is generated straight from the same markdown files the HTML pages come from, so it can never drift out of sync.
One design decision worth copying. My X posts are archived on this site, and some of them quote other people. The quoted text stays out of every machine-readable endpoint, replaced by a link and an attribution line, because other people’s writing is not mine to hand out. A dump that respects provenance is also easier for a cautious lab to accept.
robots.txt as a welcome mat
Most robots.txt files are lists of rejections. Mine now does two other jobs.
First, it explicitly allows the AI crawlers by name: GPTBot, CCBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta’s agents, PerplexityBot, Bytespider, and the rest, 23 stanzas in total. “Explicitly welcomed” is a stronger signal than “not mentioned”, and some pipelines are conservative about the difference.
Second, it starts with a comment block written for whoever, or whatever, is reading:
# Machine-readable content for LLMs and agents: # https://solmaz.io/llms.txt index of markdown mirrors of every page # https://solmaz.io/llms-full.txt the entire site as one plain-text file # Every page is also available as raw markdown by appending .md to its URL.robots.txt is the first file every crawler fetches. It might as well contain directions.
The license
This one surprised me the most. The change with the biggest payoff was legal, and it took ten minutes.
The most careful training corpora, like Common Pile, only include text with an explicit permissive license, which disqualifies roughly the entire web, since default copyright applies to everything that doesn’t say otherwise. Your beautifully written, perfectly scrapeable blog is radioactive to them unless you say the words.
So say the words. Everything on this site is now CC BY 4.0, declared in five places: a /license page, the footer of every page, a
rel="license"tag in every page head, the frontmatter of every markdown mirror, and the header ofllms-full.txt. Anyone may share, adapt, and train on my writing, as long as they say it came from me. For an immortality project, attribution is the entire point, so this trade costs me nothing.Think about the terms before copying this step. CC BY means humans can republish your writing commercially too, and once granted, the license is irrevocable for existing copies.
Cloudflare settings
If your site sits behind Cloudflare, everything above may be silently irrelevant, because since mid-2025 Cloudflare blocks AI crawlers by default for new zones. It is a fine default for people who feel scraped rather than read, and exactly wrong for what this post is trying to do. In the dashboard:
- Set AI Crawl Control to allow all AI crawlers, for both training and search.
- Make sure AI Labyrinth is off. It feeds crawlers procedurally generated garbage pages, which is a fun idea and the exact opposite of this project.
- Don’t enroll in Pay Per Crawl.
- If Bot Fight Mode is on, confirm it isn’t challenging verified bots.
You cannot fully test this from outside.
curlwith a spoofed GPTBot user agent returning 200 is a good sign, but Cloudflare verifies real crawlers by IP range, so the dashboard is the only ground truth. Cloudflare’s crawler analytics will also show you which crawlers visit and whether any got blocked, which turns the whole exercise from faith into measurement.Discoverability
Common Crawl, which feeds most open training datasets, picks what to crawl largely by how well-linked a page is. You can be perfectly scrapeable and simply never get visited. The remaining work is old-fashioned SEO with a new customer:
- The sitemap lists the markdown mirrors and both
llmsfiles alongside the HTML pages. - Register the site with Google Search Console and Bing Webmaster Tools. Several dataset pipelines bootstrap from search-engine URL lists.
- Run your important pages through the Internet Archive’s Save Page Now. Wayback-derived corpora exist, and archive.org is the most patient crawler there is.
- Inbound links remain the main lever: your GitHub profile, your repos’ READMEs, the occasional Hacker News thread. A Wikipedia citation, where genuinely warranted, outweighs everything else on this list.
To check whether any of it worked, look yourself up in the Common Crawl index. It tells you exactly which of your URLs made it into which monthly crawl.
Write something worth stealing
None of the above matters if the content isn’t worth the disk space. Modern pipelines run quality classifiers over everything they ingest, and text that reads like filler gets filtered out long before a model ever sees it. You cannot plumb your way past that, and you shouldn’t want to. The point of the exercise is to preserve thinking, so there has to be thinking.
But I keep meeting the opposite failure. People who write genuinely valuable things, deep technical explanations, hard-won practical knowledge, put them on the open web and then never spend the one afternoon it takes to make them legible to machines. Their equations render in JavaScript, their license is the default all-rights-reserved silence, and their CDN quietly serves challenge pages to every crawler. They did the hard part and skipped the easy part.
The labs still make the final call, and there is something appropriately humbling about that. You do everything right and then wait, like an author with a manuscript in the mail, except the publisher is a filtering pipeline and the acceptance letter is a model that finishes your sentences. I have done my part. See you in the weights.
- Every page on this site now has a raw markdown mirror. Append
-
People report Codex deleting their home folder or production database? Hasn't happened to me. But before someone reports their github or huggingface org being deleted: This is why you don't give your agent tokens with force-push or admin access Here is how to protect your hugging face account: (P.S. my local credential broker is almost finished and it works great on github, hf and sudo commands. Complete lockdown against agent deletion risk, without being bogged down with PRs, too many approval requests or configuration. Will launch here in a few days) -
-
-
-
No need to be offended, I'm actually a fan of your work! The metrics might have been wrong or just misfired. Looking at your recent posts, they don't have the smell Curious, as an example, was this a model, or written manually? I have the corpus here, and only some of them have the smell to me x.com/i/status/20559… -
Write-up of Reiner Pope's Lecture: How GPT, Claude, and Gemini Are Actually Trained and Served
Note: this post is an AI-assisted write-up of the blackboard lecture Reiner Pope gave on Dwarkesh Patel’s podcast. Watch the original video: How GPT, Claude, and Gemini are actually trained and served (YouTube, 2h13m).1
Pope is the CEO of the chip startup MatX and previously worked on TPU architecture at Google. With two rules of thumb (a roofline model of a GPU rack, and “set competing costs equal to each other”), he derives why batching makes tokens up to 1000x cheaper, why frontier models may be over-trained ~100x beyond Chinchilla-optimal, and how much of a lab’s serving stack you can reverse-engineer from its public API prices. The figures below are redrawn from the blackboard.
I have tried to stay faithful to the original throughout, converting the dialogue into prose and keeping all the numbers as stated. Any errors introduced in the conversion are mine.
The question that motivates everything
Dwarkesh opens with a pricing puzzle. Companies like Anthropic, OpenAI, and Cursor offer a “fast mode” that streams tokens at roughly 2.5x the speed for 6x the price. What is mechanically going on that makes this trade possible? Could you pay 100x more and go even faster? And could there be a “slow mode” where you wait minutes and pay much less?
Pope’s answer is that the dominant effect is batch size, and the rest of the lecture quantifies exactly what batching does to latency and cost. (A second effect, speculative decoding / multi-token prediction, is set aside.)
The whole analysis rests on two simplifications:
- A roofline model of the hardware. For a cluster like an NVIDIA Blackwell NVL72 rack (72 GPUs), only two numbers matter: memory bandwidth and compute throughput (FLOPs).
- Two numbers for the model. The time to operate on the weights, and the time to operate on the context (the KV cache).
The KV cache is the per-conversation state the model keeps in memory. During decode, each new token runs a full forward pass through all the weight matrices, and its attention mechanism looks back at an internal representation of every previous token. That stored representation is the KV cache, and reading it is dominated by memory fetches rather than matrix multiplies.
The two-line roofline
The time for one decode step is bounded below by whichever is slower, the memory system or the compute:
t≥max(tmem,tcompute)The compute side has to multiply a batch of B tokens by all the active parameters:
tcompute=FLOPsB⋅Nactive(The attention compute is ignored; it is small in comparison.) Note the distinction between active and total parameters: in a mixture-of-experts model like DeepSeek V3, about 37B parameters are active per token out of roughly 700B total.
The memory side has to fetch all the weights once per step, plus the KV cache of every sequence in the batch:
tmem≥memory bytes/sNtotal+B⋅lenctx⋅bytestokThese two lines are enough to draw the latency picture:
The weight fetch is a constant floor: no matter how small the batch, you must stream all total parameters from HBM into the chips once per token, and if you use all your memory bandwidth you cannot beat that. This is the latency lower bound, and it already answers the fast-mode question: for a given hardware configuration there is a floor on how fast tokens can come out, and paying more only helps until you hit it.
Cost is a different plot. Renting the GPUs for one step costs the same regardless of batch size, but the step produces B tokens, so the cost per token is t/B:
At batch size 1 the weight fetches are not amortized over anything and the economics are up to a thousand times worse. As the batch grows, the weight-fetch hyperbola vanishes and the compute term becomes a hard cost floor. This also answers the “slow mode” question: a hypothetical Claude Code Slow would live on that floor, and it would not be much cheaper than normal serving, because the compute and the KV fetches are unique to each request and cannot be amortized further.
The magic batch size
Where is the balance point where memory time equals compute time? Ignoring the KV term for a clean answer and equating the weight fetch with the weight multiply:
mem BWNtotal=FLOPsB⋅Nactive⟹B=mem BWFLOPs⋅NactiveNtotalThe first factor is purely a hardware constant. Counted in FP4 multiplies (half a byte each), it comes out around 300 on most GPUs, and it has stayed roughly stable from A100 to H100 to B100 because FLOPs and memory bandwidth grew together. The second factor is the sparsity of the model. So:
B≳300×sparsityFor DeepSeek, which activates 32 of 256 experts (sparsity 8), that gives a batch of about 2,400 sequences. In practice people run double or triple that, since real-world efficiency is worse than the roofline. Including the KV fetch would push the optimal batch higher still. Remarkably, this result depends only on sparsity, never on model scale.
Trains departing every 20 milliseconds
How does a batch fill up with real users? Pope’s model is a train schedule. The server starts a new batch every ~20 ms whether or not it is full: any requests that are ready board the train, and a request that arrives just after departure waits for the next one. Worst-case queueing latency is therefore about 40 ms.
The 20 ms itself comes from a separate design principle: you want to read your entire HBM capacity once per forward pass, so the natural step time is capacity divided by bandwidth. On the Rubin generation that is 288 GB / 20 TB/s ≈ 15 ms, and the number has hovered around 20 ms across many HBM generations. There is no point going slower, because reading the read-only weights or the KV cache twice per token does nothing for you.
A batch of ~2,000 at ~64 steps per second is ~128,000 tokens per second per rack. Google has bragged about Gemini traffic in the hundreds of millions of tokens per second worldwide, so one rack’s economical batch is about one-thousandth of Gemini. That is the economy of scale in inference: real, but reachable by any serious provider.
Does sparsity hurt quality?
The roofline says sparsity is nearly free performance, so the follow-up is empirical: how much quality do you lose? From the paper “Unified Scaling Laws for Routed Language Models”, with an older MoE technique, a 64-expert model with 370M active parameters matched a dense 1.3B model. That is a 64x increase in total parameters for a 4x effective gain, a huge parameter cost for a modest efficiency win.
And yet from the systems side it is still nearly a pure win: the extra weight fetches amortize over a larger batch, so you keep increasing sparsity until you run out of simultaneous users. The real price is memory capacity, which is what the next sections are about.
Laying out a mixture of experts on a rack
An MoE layer has a router that sends each token to a small fraction of the experts (each expert being an ordinary MLP), an all-to-all “dispatch” of tokens to their experts, an all-to-all “combine” that sums the results, and a residual connection around the whole thing.
The standard practice is expert parallelism: different experts live on different GPUs. DeepSeek’s 256 experts on a Blackwell rack (using 64 of the 72 GPUs for divisibility) means 4 experts per GPU. Since the router’s decisions are data-dependent, any GPU may need to send tokens to any other GPU.
This all-to-all traffic pattern is a perfect fit for how a rack is wired. In NVIDIA’s design the GPUs sit on the outside of the rack and NVSwitches in the middle, with every GPU cabled to every switch, so any GPU reaches any other in two hops. This is the scale-up network (NVLink). Leaving the rack means taking the scale-out network through a NIC and a data-center switch, which is typically about 8x slower.
If you spread one expert layer across two racks, half of every all-to-all crosses the slow rack-to-rack boundary and becomes the bottleneck. So one rack bounds the size of an expert layer, and this is what has been driving interconnect domains bigger: Hopper had 8 GPUs in a scale-up domain, Blackwell 72, Rubin 500-something (some of that is Jensen math, but there is a genuine ~4x from a much harder rack design). The physical constraint is mundane: cable density. Doubling the GPUs in a rack literally doubles the density of cables that must be routed to the switches, against limits of space, weight, power, cooling, and the bend radius of the cables.
This is also a lens on model scaling history. GPT-4 (2023) was rumored to be over a trillion parameters, and models only clearly exceeded that scale once racks with tens of terabytes of fast memory arrived. Google’s TPU deployments have had very large scale-up domains for a long time, which may be part of why Gemini’s pre-training scaled successfully early. The summary: active parameters are limited by compute cost, and total parameters are limited by scale-up size.
Pipeline parallelism
Expert parallelism uses up one rack. To use more racks, the remaining options are data parallelism and pipeline parallelism (tensor parallelism has become irrelevant now that experts are small). Pipelining means putting different layers on different racks: a token flows through rack 0 for the first stage of layers, then hops to rack 1, and so on.
Is the hop a bottleneck? Compare the time spent on scale-up traffic to the time on scale-out traffic. Crossing racks sends each token once per stage, while inside the rack each token fans out to every activated expert, twice (dispatch and combine), for every layer in the stage:
tscale-outtscale-up=81⋅2⋅(activated experts)⋅(layers per stage)≥1The 1/8 is the bandwidth ratio. With 8+ activated experts and a few layers per stage, the inequality is easily satisfied, so an entire pipeline of racks, one stage each, is communication-feasible.
Dwarkesh brings up Ilya’s remark that “as we now know, pipelining is not wise,” and the architectural constraints it imposes (e.g. Kimi’s attention to layers a few back is awkward to pipeline). Pope’s framing: pipelining is a massive hassle with real but narrow benefits. It saves no runtime at all (the memory fetches just happen on a different rack), but it divides the weight storage per rack, which matters if memory capacity is your constraint.
The catch is micro-batching. To keep four racks busy, you need four micro-batches in flight, each wrapping around for its next decode step as soon as it finishes:
In inference this is natural and the bubble costs nothing; latency is identical to running unpipelined on one rack. In training there is a hard stop between the forward and backward passes of a batch, which creates a genuine bubble of idle time (the literature has zero-bubble and one-forward-one-backward schemes to interleave around it; as Dwarkesh notes, you could also mine Bitcoin in it):
The training batch size itself is a trade-off: smaller batches are always better for ML convergence (fresher gradients), larger batches are better for systems throughput, and the optimum sits in between.
The memory wall and why the KV cache won’t shard
Here Dwarkesh raises the macro puzzle. Memory is the scarce commodity of the moment: Dylan Patel claims hyperscalers are spending half of their CapEx on memory, and consumer devices are getting squeezed. Yet the pipelining analysis just said racks have a memory surplus, since a trillion-parameter model needs only ~1 TB against a rack’s tens of TB. Why is Jensen shoving all that HBM in?
Write down the memory demand across the whole system:
Cmem=Ntotal+B⋅lenctx⋅bytestokSharding across E GPUs of expert parallelism and P racks of pipelining, the per-GPU requirement is this divided by E⋅P. But the global batch is (number of micro-batches) × (micro-batch size), and the number of micro-batches needed to fill the pipeline equals P, while the micro-batch size b is pinned near 300×sparsity by the roofline. Substituting B=P⋅b, the P‘s cancel in the KV term:
cmemper-GPU=E⋅PNtotal+Eb⋅lenctx⋅bytestokMore pipeline stages keep shrinking the weight footprint, but the KV footprint per GPU stays constant: each extra stage requires proportionally more sequences in flight to stay busy. The KV cache can’t be amortized across the batch (it is unique per user), and it can’t be sharded across pipeline stages either. It loses on both fronts, and once you pipeline even a little, it becomes the dominant use of memory.
So what do labs actually run? Per the DeepSeek paper: expert parallelism up to the scale-up domain size, then very little pipelining (maybe none, maybe 2 stages so the weights aren’t an issue). Frontier inference essentially lives inside a single scale-up domain. Each rack hop would also add on the order of milliseconds of latency per token, which stacks across stages in sequential decode.
The last piece of the scale-up story is bandwidth. The weight-fetch latency is
tmem, weights=S×BW per GPUNtotalwhere S is the scale-up size, because all GPUs in the domain load the weights in parallel. Per-GPU HBM bandwidth improves maybe 1.5–2x per generation, but S jumped 8x from Hopper to Blackwell. Pipelining solves the capacity problem; big scale-up domains solve the bandwidth problem, which is what actually lets you serve at low latency and long context.
Over-trained 100x beyond Chinchilla
Chinchilla scaling tells you the compute-optimal ratio of model size to training data. But a lab does not minimize training compute; it minimizes total compute across pre-training, RL, and inference for all its users. Pope’s heuristic: when minimizing a sum of competing costs, the minimum tends to sit where the costs are equal (true for x+1/x, for ex+e−x, and generally for power laws). So set all three equal.
Using the 6ND rule (6 FLOPs per parameter per token for forward+backward, 2 for forward only):
ctotal=pre-training6NactDPT+RL[2 to 6]NactDRL⋅inefficiency+inference2NactDinfRL sits between 2 and 6 because you generate every rollout but may not train on all of it, and it carries an extra inefficiency factor (~30%) because RL involves a lot of decode, which runs at lower MFU than training. The active parameter count divides out entirely. Working through the arithmetic on the board, the equal-cost condition lands at roughly
DPT≈1.5DRL≈DinfIn words: the number of pre-training tokens, RL tokens, and lifetime inference tokens should all be about the same, within factors the analysis can’t resolve. (Dwarkesh’s gloss: every model should stream out roughly the sum of human knowledge that was streamed into it.)
Now plug in real-world guesses. Global traffic of ~500M tokens/s, cut by 5–10x for one specific model in a family, gives ~50M tokens/s; times a two-month deployment life, that is roughly 2.6×1014, call it 200T inference tokens. The rumor mill says frontier pre-training is ~150T tokens, which matches. With ~100B active parameters, Chinchilla would prescribe only ~2T tokens. The ratio is about 100x over-trained, derived almost from first principles. As Pope puts it, approximate everywhere, set A equal to B, and it’s kind of empowering how far that gets you.
One asymmetry he flags: if your model might miss the frontier and get thrown away, the expected inference tokens shrink, so you should derate the inference term and err toward less over-training.
Reading the infrastructure off API prices
Since providers are incentivized to price close to cost (otherwise someone scoops them), public price sheets leak infrastructure details.
The 200k context surcharge
Gemini 3.1 charges 50% more per token beyond 200k context. Redraw the roofline as a function of context length at a fixed large batch: compute cost is flat (the attention FLOPs slope is negligible until millions of tokens), while memory cost grows linearly with the KV cache. The provider wants to be profitable at every context length, so a two-tier price is laid over a kinked cost curve, and the price bump should sit near the crossover where the model flips from compute-bound to memory-bound:
Assuming the crossover is at 200k, you can solve for the model’s KV bytes per token. Setting KV fetch time equal to compute time and cancelling the batch size:
bytestok=FLOPsmem BW⋅lenctxNact=3001⋅200k100B≈1.7 kB per tokenIs ~2 kB/token plausible? The KV size is (number of unique attention contexts) × 2 × dhead × (KV heads). With dhead=128 and 8 KV heads and a single global context shared across all layers (the Character AI trick, also used in Gemma), you get exactly 2 kB. Sparse attention gets there with bigger raw numbers divided by the sparsity. So the pricing page is consistent with a real architecture, if maybe a little on the small side.
Input vs output prices
Output tokens cost 3–5x more than input tokens. The two phases differ in tokens per forward pass: decode processes one new token per pass, prefill processes the whole prompt in one pass. Dividing the same roofline by the tokens per pass (
len_pass) gives the cost per token: the compute term is flat, and the memory term is a hyperbola that only bites whenlen_passis small.Prefill is compute-bound, decode is memory-bandwidth-bound, and a 5x price gap says decode at the provider’s operating point is deeply memory-bound: they are paying ~5x more per output token in memory time than the compute floor.
This also explains the context length plateau. Contexts jumped from ~8k (GPT-3 era) to 100–200k around GPT-4 and have hovered there for a year or two, which suggests that is the balanced cost point. The barrier to 100M-token contexts (the “in-context learning is enough for AGI” scenario) is memory bandwidth and capacity, and HBM is not getting hugely better. Sparse attention (DeepSeek published one mechanism that effectively puts a square root on the KV term) is a big one-time improvement, but going too sparse costs quality, so it is a get-out, not a solution.
Cache pricing and memory tiers
Providers charge much less for cached input tokens, and charge different rates for keeping a cache alive 5 minutes vs 1 hour. There are two ways to produce a KV cache for a token: rematerialize it from scratch (a forward pass: pure compute cost, nothing to store) or hold it in some memory tier (near-zero retrieval cost, but you occupy capacity that scales with hold time). Each tier down (HBM → host DDR → flash → spinning disk) is cheaper to occupy and slower to retrieve from.
Which tier backs which price? Pope’s rule: a storage tier is well-matched to hold times around its drain time, capacity divided by bandwidth, the same ratio that gave 20 ms for HBM. DDR drains in seconds, flash in about a minute, spinning disk in about an hour. So a 5-minute cache tier and a 1-hour cache tier probably map to flash and spinning disk, which surprised him: “I’m kind of shocked to see spinning disk being used at all.”
Convergent evolution with cryptography
The sit-down portion covers Pope’s blog post on how neural nets and ciphers evolved similar shapes. Both need to thoroughly mix information across all their inputs, and even stirring cake batter alternates directions for the same reason. But they optimize in opposite directions. A neural net is kept differentiable in a useful way: residual connections and LayerNorm exist to keep the derivative simple and meaningful for gradient descent. A cipher is designed so that its derivative is useless: differential cryptanalysis attacks a cipher by differentiating it (over the field of two elements), and a well-designed cipher makes a small input difference blow up into a huge output difference, the avalanche effect. Adversarial examples in image models are exactly the avalanche property showing up where it is not wanted.
Building ciphers out of neural nets is a bad idea (99% of new ciphers get broken), but one construction has productively flowed the other way. A Feistel network turns any non-invertible function f into an invertible two-input block:
g(x,y)=(y+f(x),x)To invert, read off x from the second slot, then recover y=z−f(x).
The 2017 RevNets paper imported this into deep learning: make each layer a Feistel block (which turns out to look like a residual connection from two layers back) and the whole network becomes invertible. Training normally has to write every layer’s activations to HBM on the forward pass so that the backward pass can read them, a memory footprint linear in depth and often the largest one in training. An invertible network stores none of it: during the backward pass it undoes the forward pass in lockstep, rematerializing activations as needed.
That is spending compute to save memory, the exact mirror image of the KV cache, which spends memory to save compute. Given where hardware is, the KV cache direction is usually the profitable one, which is a fitting last word for a lecture that is mostly about the price of memory.
-
How this post was made, in the interest of transparency. I first had one agent (Codex) prepare the raw material. My prompt, verbatim:
https://www.youtube.com/watch?v=xmkSf5IS-zw
download using yt dlp. create a jpeg screenshot every 30 seconds and save it
also get the transcription for it from dwarkesh’s website if it exists. if not, you can transcribe using the whisper model in bob@isengard
save it in a folder. I want to prepare it for an another agent to prepare it for further processing
That produced the video, 268 frames at 30-second intervals, and Dwarkesh’s published transcript (no Whisper needed). I then handed the folder to a second agent (Fable, in Cursor), which read the transcript, inspected the frames, and redrew the blackboard diagrams as matplotlib figures. My prompt, verbatim:
convert this to a write-up that is faithful to the original. use the video and images if needed. create figures based on what is on the screen and such
The draft first lived as a standalone GitHub document (“just make it standalone, dont add it to my blog”, then “create doc in ~/scratch repo, github markdown. i will read that. make sure it renders nicely”) before I asked for this post (“ok I want you to create a blog post citing the dwarkesh podcast in my blog, linking the youtube and just saying that this is a write-up of the video”). In between I reviewed the output and asked for fixes, for example on the cost figure (“the graph after this: the lines are a bit tight. could we make it clearer?”) and on figure rendering (“also, in the svgs, the space between some things are too much. if you can improve the text rendering in the svgs, it would be great”). ↩
-
Theoretical Upper Bounds for LLM Performance
Note: this post is AI generated, adapted from my working notes behind the Local Frontier calculator. It is a work in progress and may lack rigor in places. If you find an issue in the formulation, please email me at [email protected].
How fast can a machine serve a large language model? I built Local Frontier, an interactive database of local AI hardware and model profiles, to answer that question for any hardware and model pair. This post derives, from first principles, the math the calculator runs. That math is a roofline, an upper bound that says what a given hardware and model pair cannot exceed under a stated set of assumptions. Nothing in it is specific to local machines. The bounds are built from memory capacity and bandwidth alone, so the same formulas cover a MacBook and a datacenter GPU node. The post focuses on local hardware because that is what Local Frontier compares. Its outputs are ceilings for comparing machines, and a real implementation lands somewhere below them.
The argument builds in layers, each created by a problem the previous layer cannot solve. A memory system is two numbers, capacity and bandwidth, and their product turns out to be a natural measure of what the system can fund. That product yields a clean but loose throughput ceiling for batched decoding. The loose ceiling ignores per-session context traffic, so it is too generous for real serving, and repairing it gives the bound the calculator actually uses. The repaired bound is then filled in per architecture by small adapters and checked against real hardware.
Interactive figures accompany the derivation. They all use the same toy setup, a 32B-class dense model at 4-bit (18 GB of weights, 6 GB of runtime overhead, 0.26 GB of KV per thousand tokens of context) on a 128 GB machine with 800 GB/s of sustained bandwidth, so the numbers stay comparable from figure to figure.
Three questions
Serving an LLM raises three separate questions, and mixing them up is an easy way to be wrong about a machine.
Question What it asks Resident fit Can the model plus runtime overhead be held in memory at all? Single-session speed What is the memory-side ceiling for one active conversation? Useful serving throughput Across many active sessions, how many tokens per second can the device produce while each session stays above a minimum useful rate? The third question is the hard one. A machine can fit many sessions in memory and still be too slow per session at that concurrency. So the calculator must report both whether sessions fit and whether the fitting sessions are fast enough to be worth running.
Every number produced here is an upper bound. A real system can fall below it for reasons the memory model deliberately ignores, among them compute limits, kernel quality, quantization overhead, scheduling, CPU involvement, paging, interconnects, and thermal throttling. The value of a clean upper bound is that it tells you the best case you are allowed to hope for, and therefore how much room an implementation still has.
Memory power
Before any model-specific detail, ask what a memory system fundamentally offers. When people compare accelerators for local inference they list many specs, from capacity and bandwidth through compute throughput, cache hierarchy, PCIe lanes, and thermals. For the decode phase of autoregressive generation, two of these recur in almost every bound. How much state can the memory hold, and how fast can it move that state? We start there.
Two numbers and their product
Model an idealized memory system by exactly two quantities.
C=usable memory capacity,R=sustained memory bandwidth.Capacity is a stock, the amount of resident state that can exist at once. Bandwidth is a flow, the amount of state that can be moved per second. They have different units and answer different questions, so any single product of them needs justifying before we rely on it.
Define memory power as
D=CR.If C is in GB and R is in GB/s, then D is in GB2/s. Power is meant in the colloquial sense of capability, as in computing power, and no watts appear anywhere in this model. The units look strange, so the rest of this section explains why this particular combination of C and R is the right scalar.
For a feel of the scale, a 24 GB GPU at 1000 GB/s has D=24×1000=24,000 GB2/s. One thing to be careful with throughout is that every memory quantity in a given calculation must use the same unit system. The equations are identical in bytes, GB, or bits, and only the numeric value of D changes.
Feasibility theorem
Here is the toy problem that justifies CR. The models in this formulation are deliberately simple, and they are meant as rules of thumb, useful when deciding on hardware for a model or when checking whether an existing setup is getting the most out of the hardware it runs on.
A pure memory workload is a pair w=(h,r), where h is the resident information it must keep alive and r is the information flow rate it must sustain. The workload is feasible on the system M(C,R) when the memory can both hold the state and carry the traffic. There are exactly two ways to fail.
The first is a capacity failure. If a workload needs more resident state than the memory can store, it cannot fit, and no scheduling trick repairs that.
h≤C.The second is a flux failure. If a workload needs more traffic per second than the interface can deliver, it cannot be sustained, and no surplus capacity repairs that.
r≤R.In the idealized model these two conditions are also sufficient. If h≤C and r≤R, allocate h of state and stream at rate r. Therefore the feasible set is precisely the rectangle
F(C,R)={(h,r):0≤h≤C, 0≤r≤R},whose area is
μ(F)=CR.So memory power has a concrete meaning. It is the measure of the feasible workload region. A workload lives at a point (h,r) in the plane, the system can serve every workload inside its rectangle and none outside, and the size of that rectangle is CR.
A caution before building on this. The rectangle is a toy model, and the world is messier in both directions. Real hardware delivers only a fraction of its catalog bandwidth. Even a perfectly sequential read falls short of the spec number, and the achieved fraction depends on the access pattern, on whether enough parallel work is in flight to hide memory latency, and on the hardware itself. The sufficiency claim is idealized too, because the rate a machine reaches depends on which bytes are read and in what order. The bounds below survive this, since real inefficiency only pushes a machine further below its ceiling. The cost lands on comparisons between machines. If one machine sustains 85% of its spec bandwidth and another 60%, ceilings computed from spec numbers make the second machine look better than it really is. Read R as sustained bandwidth where a measured number exists, and read every comparison in this post as approximate.
A caveat on the metric
One misreading is worth heading off. The product CR measures the workload-feasibility region, the set of jobs the device can support at an instant. A different question, “how many distinct memory histories can this device produce over a time T”, has a different answer, namely the resident information plus the information streamed over that window,
C+RT.The feasibility-region reading is the one an inference-throughput bound will need, because serving means keeping sessions resident in memory and streaming data for them.
A first decode bound
The feasibility theorem describes static workloads. Decoding is dynamic, emitting tokens over time. This section turns memory power into a throughput ceiling for batched decoding, and in doing so shows where the product CR enters a real performance bound.
A toy decoder
Consider the decode phase of an autoregressive transformer, simplified to the memory system alone. Let
W=model weight footprint,K=per-session KV/state footprint,let b be the batch size (number of concurrent sequences), and let s be the number of decode steps per second. Each decode step emits one token per active sequence, so aggregate output throughput is
T=bs.Throughput is a product of two things, how many sequences run in parallel and how fast the shared model can be swept, and the memory system caps each factor separately.
Capacity limit
The model weights must be resident once, and every active sequence needs its own KV cache, the per-conversation attention state that grows with context length. So the resident state is W+bK, and it must fit.
W+bK≤C⟹b≤KC−W.This is the capacity limit, and it sets the maximum parallelism. It is exactly the h≤C condition of the feasibility theorem, with h=W+bK.
Bandwidth limit
In a dense transformer, each decode step must apply the model weights. In the memory-bound idealization, applying the weights means streaming roughly W of data per step. With bandwidth R, the step rate obeys
sW≤R⟹s≤WR.This is the bandwidth limit, and it sets the maximum step rate. It is the r≤R condition, with the per-step traffic playing the role of the flux.
Memory-power decode bound
Multiply the two caps. Throughput is parallelism times step rate, and each is separately bounded, so
T=bs≤KC−W⋅WR=KWR(C−W).Substituting D=CR and factoring out C exposes the memory-power term.
T≤KWD(1−CW).When the model is much smaller than memory, W≪C, the correction vanishes and the bound collapses to the memorable form
T≲KWD.This is the memory-power decode bound. The product CR appears because maximum throughput genuinely factors into maximum parallelism times maximum step rate.
how many sessions residentKC×how many sweeps per secondWR=KWCR=KWD.The numerator D=CR is the machine. The denominator KW is the workload, the per-session state times the model sweep size. The bound reads as throughput is memory power divided by memory cost per active model-token.
Scope of the bound
The memory-power decode bound governs batched throughput. Set b=1 and it degenerates to
T≲WR,so for a single session only bandwidth matters and capacity is merely a fit constraint. That is the correct behavior, and it makes clear what each metric is for.
Use case The metric that governs it Single-user local chat bandwidth R, with capacity C as a fit gate Largest model that fits capacity C first, then bandwidth Maximum batched decode throughput memory power D=CR Memory power is the right scalar precisely for the third row. The next section explains why even that bound is too optimistic.
Missing traffic
The memory-power decode bound assumes the only per-token memory traffic worth counting is the model sweep W, shared across the batch. Real decoding also reads each session’s growing KV cache, and that traffic is private, so it does not amortize over the batch. Ignoring it makes the bound promise throughput that long-context serving can never reach. This section introduces the correct per-token accounting and shows the memory-power bound falls out of it as a loose corollary.
Universal resource bound
Step back to the most general statement, which holds regardless of architecture. If each output token costs at least qmin of unavoidable memory traffic and at least amin of unavoidable compute, and the device delivers at most R bytes/s and F FLOP/s, then over all setups that fit,
Tmax≤setup fitsmaxmin(qminR, aminF).Two rooflines, and the workload lives under the lower of them. Throughout this post I assume decode is memory-bound, which is the usual case for local serving, and keep the memory roofline,
Tmax≤qR,where q is the memory traffic per output token. Everything now reduces to estimating q honestly. The focus on the memory side is also practical. A hardware catalog can collect capacity and bandwidth consistently across consumer and workstation devices, while comparable sustained-compute numbers are much harder to obtain.
Bytes per token
Split the per-token traffic into the two kinds that behave differently under batching. The model (or active-expert) weights are shared, since one sweep serves the whole batch, so their per-token cost is divided by the batch. The KV/context read is private, since each session reads its own cache, so its per-token cost is not divided at all. With Wactive the shared weight traffic per iteration, ρ the tokens emitted per session per iteration (one for ordinary decoding), and Kread(L) the private context traffic per output token at active context length L,
qsimple(b)=bρWactive+Kread(L).The first term shrinks as the batch grows, because more sessions share each weight sweep. The second term does not move. A larger batch does not make any session’s context cheaper to read, and this asymmetry is the main reason long-context serving behaves differently from short-context serving.
Interactive figure: bytes per output token as the batch and the read context change. Enable JavaScript to explore it.
Memory-power bound as a corollary
Treat qsimple as an optimistic lower estimate of the true bytes per token, qsimple(b)≤qactual(b). Dividing R by a smaller denominator gives a larger quotient, so substituting qsimple keeps the result an upper bound.
Tmax≤b:Wresident+bKstore+O≤Cmax bρWactive+Kread(L)R.Here the maximization runs over batches that fit in memory, with Wresident the resident model footprint, Kstore the KV memory stored per session, and O the runtime overhead. This is the honest simple bound, bandwidth divided by shared-per-token cost plus private-per-token cost, maximized over fitting batches.
Now recover the previous section’s bound by deliberately throwing information away. Since Kread(L)≥0, dropping it only loosens the denominator.
bρWactive+Kread(L)R≤bρWactiveR=WactivebρR.The capacity constraint caps the batch at b≤(C−Wresident−O)/Kstore, so
Tmax≤ρKstoreWactiveR(C−Wresident−O)=ρKstoreWactiveD(1−CWresident+O).This is the memory-power decode bound again. It is what you get by discarding the private context term and using the largest batch that fits. That is why it can sit far above achievable throughput while remaining a true ceiling. The ordering is
actual throughput≤simple (KV-aware) bound≤memory-power bound.The memory-power bound is the orientation line. The KV-aware bound is the one to serve from, and the next section develops it into the operational calculator.
KV-aware bound
This section turns the simple bytes-per-token bound into the model the calculator runs. Three refinements are needed. Context must be split into the part that controls memory and the part that controls speed. The shared weight traffic must be allowed to grow with the batch, which matters for mixture-of-experts (MoE) models, models that route each token through a small subset of their weights. And the batch must be filtered so that we never count concurrency at which every session has become uselessly slow.
Two context lengths
A serving system usually reserves KV space for a long maximum context but reads, on average, a shorter active context. These two lengths drive different parts of the bound, so we keep them separate.
LallocLreadLread=reserved (maximum) context,=average active context,≤Lalloc.The allocation length controls how much KV memory each session reserves, and therefore how many sessions fit. The read length controls how much context each output token must stream, and therefore per-token cost. Collapsing them into one number either overcharges memory or overcharges speed.
Model quantities
A model contributes five quantities. Two are about fitting, two are about speed, and one is about decoding style.
Symbol Role Wresident Full resident footprint (for MoE, all resident weights, including the inactive experts) Wbatch(b) Shared weight traffic per decode iteration at batch b Kalloc(Lalloc) KV/cache memory reserved per session, which controls concurrency Kread(Lread) Private context traffic per output token, which controls decode cost ρ Tokens emitted per session per iteration (ρ=1 ordinary, ρ>1 speculative) The split between Kalloc and Kread mirrors the split between the two context lengths. Allocation controls how many sessions fit, read controls how fast each one decodes. Note that Wbatch now depends on the batch size b, for a reason the adapter section explains.
Memory-fit batch
The first gate is whether sessions fit. Load the model, reserve overhead, and divide the remainder by the per-session allocation.
bmem(Lalloc)=⌊Kalloc(Lalloc)C−Wresident−O⌋.This is memory-fit concurrency only. It is necessary but not sufficient, and it is exactly the trap that makes a machine look like it can serve a hundred sessions when it cannot serve them usefully.
Interactive figure: how many sessions fit in memory as capacity and reserved context change. Enable JavaScript to explore it.
Aggregate and per-session ceilings
The per-token traffic is the shared weight sweep amortized over the emitted tokens, plus the private context read.
qKV(b,Lread)=bρWbatch(b)+Kread(Lread).Then the memory roofline T≤R/q gives the aggregate ceiling at batch b,
T(b,Lread)≤qKV(b,Lread)R,and dividing by the batch gives the per-session rate,
r(b,Lread)=bT(b,Lread)=Wbatch(b)+bρKread(Lread)ρR.As b grows, the aggregate T rises but the per-session r falls. That tension is the whole serving tradeoff, and it is why a fit-only bound is not enough.
Usable-batch correction
The fix is to refuse batches at which a session would crawl. Impose a per-session floor r⋆, the minimum useful tokens/s/session, and solve r(b)≥r⋆ for b. Replacing the batch-dependent Wbatch(b) by its shared lower bound Wactive keeps a closed form. Because Wbatch(b)≥Wactive, the substitution only weakens the condition, so the implication runs one way, and the closed form is a necessary condition on the admissible batch rather than a sufficient one.
r(b)≥r⋆⟹b≤ρKread(Lread)ρR/r⋆−Wactive,which defines a rate-limited batch
brate(Lread,r⋆)=⌊ρKread(Lread)ρR/r⋆−Wactive⌋.The usable batch is whichever gate binds first.
busable=min(bmem(Lalloc), brate(Lread,r⋆)).Because brate comes from a necessary condition, busable is itself an upper bound on the truly admissible batch, and the calculator applies the exact floor test with the true Wbatch(b) in the next step. This is what stops the “hundred sessions” illusion. As context grows, Kread(L) grows, so brate falls quickly even while bmem stays large. The KV slots fit while the useful rate does not.
Interactive figure: aggregate and per-session ceilings against batch size, with the memory gate and the per-session floor. Enable JavaScript to explore it.
The bound the calculator uses
Collecting the pieces, define the usable batch set as the fitting batches that also clear the floor,
B(Lalloc,Lread,r⋆)={b:1≤b≤bmem(Lalloc), b1qKV(b,Lread)R≥r⋆},and take the best aggregate over that set.
Tmax(Lalloc,Lread,r⋆)≤b∈Bmax bρWbatch(b)+Kread(Lread)R.This is the KV-aware bound, the main practical formulation. In words, try every batch that fits, reject the ones too slow per session, and for the rest take bandwidth divided by bytes per output token, keeping the best.
The looser memory-power bound is its corollary, obtained as before by dropping the private term and using the largest fitting batch.
Tmax≤ρKalloc(Lalloc)WactiveD(1−CWresident+O),D=CR.The two stand in a fixed relation, which is the main result of the derivation, written first by name and then in full.
TmaxTmax≤KV-aware bound≤memory-power bound≤large-memory limit≤b∈BmaxbρWbatch(b)+Kread(Lread)R≤ρKalloc(Lalloc)WactiveD(1−CWresident+O)≤ρKalloc(Lalloc)WactiveDThe gap across these terms is the point of the whole derivation. The KV-aware line is the tight, practical bound. The memory-power line shows the memory system’s large theoretical capacity-bandwidth product, and the distance between them comes from private context traffic, expert diversity, and the per-session floor. The final term drops the resident-model factor as well, so the right-hand side is exactly the simplified D/(KW) from the memory-power decode bound, now with K=Kalloc(Lalloc) and W=Wactive. It is the loosest, most optimistic reading, since the resident model and overhead always claim a real share of C.
Single session
Set b=1 to recover the latency-style bound for one conversation. There is no batch to amortize the weight sweep over.
T1,max≤Wbatch(1)/ρ+Kread(Lread)R,and for ordinary decoding (ρ=1) this is just R/(Wbatch(1)+Kread(Lread)). Capacity has dropped out except as the gate that decides whether the model fits at all, consistent with the earlier observation that bandwidth governs single-session speed.
Speculative decoding
Speculative decoding lets ρ>1. A draft model proposes several tokens and the target verifies them in one iteration, so ρ is the expected number of accepted tokens per session per verification step, bounded by the draft length γ as 1≤ρ≤γ+1. The temptation is to multiply throughput by ρ and stop. That is wrong, because the draft model and verification are not free. Their traffic belongs in Wbatch(b) or in the per-token term. The safe rule is to never scale by ρ without charging the draft cost in the denominator. With both effects included, speculative decoding moves through the same KV-aware formula unchanged.
The two bounds side by side
The gap between the memory-power bound and the KV-aware bound is easiest to see as memory traffic. Below, two copies of the same machine decode side by side. Each board is the machine’s usable memory, with the weights packed into an orange container of equal-sized cells and each session’s blue KV cells packed into a small container of its own.
Each decode iteration must move every byte that its accounting charges, at the same bandwidth on both machines, so the charged cells light up one by one and the board resets when the iteration completes. The left board charges only the shared weight cells, so it resets quickly and its token counter races ahead. The right board also charges the read cells of every session’s context, so its iterations stretch as the batch and the context grow. A real iteration takes milliseconds, so time runs in slow motion here.
Interactive animation: memory-power accounting and KV-aware accounting decoding side by side on the same machine. Enable JavaScript to watch it.
Both boards run on the same silicon at the same bandwidth, and only the bookkeeping differs between them. The left counter is the memory-power accounting at the chosen batch, and the right one is what the KV-aware bound admits once private context reads are charged.
The sliders show the two context lengths at work. The reserved context sets how many cells each session’s container holds and can push the machine past its capacity, so raising it eventually makes the containers stop fitting. The read context sets how many of those cells light up every iteration, and the longer it gets the smaller the weights’ share of each iteration becomes. It follows the reservation at an adjustable fraction, 90 percent by default, with the unread rest capped at 32k, and the sliders for all of this sit under Advanced. Growing the model itself slows both boards down in step, while the orange container eats the room the blue ones need.
The full memory-power bound goes one step further than the left board. It grows the batch until memory is completely full of KV cells, which is exactly what the default maximize batch mode does, and the line under the left board reports that number. Shrinking the reserved context makes it explode.
At a 4k reservation about a hundred sessions fit and the bound climbs past 4,000 tok/s on this toy machine, and at a 1k reservation it would pass 17,000. Those numbers are true ceilings and useless forecasts at the same time. What stops a real machine long before then is reading each session’s context, which is exactly the traffic the right board charges.
Model adapters
The KV-aware bound is architecture-agnostic, and an architecture enters only through the five quantities. So each model family is captured by a small adapter that supplies them.
Adapter(M)=[Wresident, Wbatch(b), Kalloc(Lalloc), Kread(Lread), ρ].Three adapters cover the catalog. They handle dense transformers, mixture-of-experts, and hybrid/sliding/recurrent attention.
Dense transformers
For a dense model with Ptotal parameters at ew bytes each, all weights are touched every step, so the resident footprint and the per-step sweep coincide.
Wresident=Wbatch(b)=Ptotalew.With Nlayers layers, Nkv key/value heads, head dimension dh, and KV byte widths eKV,store and eKV,read, full-context attention reserves and reads
Kalloc(L)=2NlayersNkvdheKV,storeL,Kread(L)≈2NlayersNkvdheKV,readL,where the factor 2 counts keys and values. Weight precision and KV precision are independent settings. NVFP4 weights (NVIDIA’s 4-bit floating-point format) do not imply an NVFP4 cache, so ew and eKV are tracked separately.
Mixture-of-experts
An MoE model is where the constant-Wbatch assumption breaks, and fixing it is the single most important adapter correction. Let Ptotal be total parameters, Pactive the active parameters per token, E the number of routed experts, and k the experts selected per token. Assuming uniformly sized routed experts, each routed expert holds
pexpert=E−kPtotal−Pactive,and the always-on remainder (dense trunk, shared experts, embeddings, attention) is
Pfixed=Pactive−kpexpert.The naive model assumes a batch touches the same active experts every session, keeping Wbatch constant. That is false. Independent sessions route to different experts, so a larger batch touches more distinct experts. With bρ token-routings, each independently missing a given expert with probability 1−k/E, the expected number of distinct experts touched is
m(bρ)=E(1−(1−Ek)bρ),and the per-iteration shared traffic is the fixed part plus the touched experts.
Wbatch(b)=ew[Pfixed+pexpertm(bρ)].At bρ=1 this reduces to the active-parameter footprint, and as bρ→∞ it saturates at all E experts. This rising Wbatch(b) is why MoE batching does not amortize for free, and why the MoE rows in the worked table reach their throughput optimum at modest batch sizes.
One caveat applies here. Every other traffic term in the bound is a deliberate under-estimate of real traffic, which is what makes R/q a true ceiling. The expert count m(bρ) is different. It is an expectation under independent, uniform routing rather than a lower bound. Real routing is correlated, since load-balancing losses push toward uniform while hot experts and topically similar sessions pull the other way, and correlated routing touches fewer distinct experts than the formula predicts. In that case the modeled traffic overstates the actual traffic, and the computed ceiling can sit below the true one. When a guaranteed ceiling is required, replace m(bρ) by its minimum k, which replaces Wbatch(b) by Wactive. The expectation form is the better estimate, the floor form is the safe bound.
Hybrid, sliding, and recurrent attention
Models with local or sliding-window attention, compressed or latent attention, or linear/recurrent state must not use the full-KV formula blindly, because their cache does not grow linearly in L everywhere. Split both KV terms into global, local, and fixed-state parts.
K∙(L)=Kglobal,∙(L)+Klocal,∙(L)+Kstate,∙,∙∈{alloc,read}.A simple read approximation with sliding-window width w is
Kread(L)=κglobalL+κlocalmin(L,w)+Kstate.Full-attention layers pay for the whole context, sliding-window layers pay only up to the window, and recurrent or latent state adds a fixed or slowly growing term. The same shape covers Gemma-style local/global attention and DeepSeek-style compressed/sparse attention, with only the coefficients changing.
Calculator procedure
The bound is now ready to compute. Because the same computation runs for every hardware-and-model pair, it is worth stating once as a procedure.
The inputs are a hardware row (C,R,O), a model adapter (Wresident,Wbatch(⋅),Kalloc(⋅),Kread(⋅),ρ), and workload assumptions (Lalloc,Lread,r⋆).
- Compute the resident margin C−Wresident−O. If it is negative, the model does not fit, so stop.
- Compute Kalloc(Lalloc) and the memory-fit batch bmem.
- For each integer batch 1≤b≤bmem, compute Wbatch(b) and qKV(b,Lread).
- Compute the aggregate ceiling R/qKV and the per-session ceiling R/(b,qKV) at each batch.
- Keep the batches whose per-session ceiling is at least r⋆.
- Among the kept batches, choose the one with the largest aggregate ceiling, and report it as the KV-aware result with its batch as b⋆.
- Separately compute the memory-power ceiling for orientation.
The output is a stack of gates, and the right phrasing depends on which gate bound.
State Meaning Resident fit The model plus overhead fits in memory Session fit At least one reserved-context session fits Floor fit Some fitting batch clears r⋆ No floor Sessions fit, but no batch clears r⋆ The common invalid reading is that fitting in memory implies serving usefully. A model can pass resident fit and session fit and still have an empty usable batch set, because every fitting batch is below the floor. The honest report for that case is “fits, but no batch satisfies the floor”, a distinct verdict from a true fit failure. Keeping the two apart is the reason the floor gate exists.
Worked examples
Now check the theory against real hardware. Consider two 128 GB machines, one bandwidth-rich and one bandwidth-poor, which isolate the effect of R at fixed C.
Two machines
NVIDIA’s DGX Spark, a small desktop AI machine, carries 128 GB of LPDDR5x unified memory at 273 GB/s. An Apple M5 Max with a 40-core GPU reaches 614 GB/s and is configurable to 128 GB of unified memory. At equal capacity their memory powers are
Hardware C R D=CR DGX Spark 128 GB 273 GB/s 34,944 GB2/s Apple M5 Max 128GB 128 GB 614 GB/s 78,592 GB2/s Both bandwidth numbers are catalog figures rather than measured sustained rates, so the earlier caveat applies. If one machine sustains a larger share of its spec than the other, the comparison will make the other machine look better than it really is.
Three models
Three MoE models, modeled from their published cards as adapter parameters, without re-measurement.
- Qwen3.6-35B-A3B, with 35B total / 3B active parameters, E=256 routed experts, k=8 routed (plus one shared) per token, and weights quantized NVFP4.
- Gemma 4 26B-A4B-it, with 26B total / 4B active, E=128, top-8 routing, hybrid local/global attention, and NVFP4 weights.
- DeepSeek V4 Flash (DS4), with 284B total / 13B active, E=256 routed plus one shared, k=6 per token, million-token context via compressed/sparse attention, and weights at a Q2-style mixed quantization.
All rows use the Local Frontier defaults, namely reserved context Lalloc=100,000, active context Lread=32,000, per-session floor r⋆=20 tok/s/session, ordinary decoding ρ=1, runtime overhead O=8 GB, and the memory roofline only. The numbers are memory-side upper bounds from the simplified adapters, to be read as ceilings for comparing hardware.
Results
Hardware Model Single-session b⋆ KV-aware aggregate Memory-power ceiling DGX Spark Qwen3.6-35B-A3B ≤ 149 tok/s 17 ≤ 345 tok/s ≤ 18.2k tok/s DGX Spark Gemma 4 26B-A4B-it ≤ 120 tok/s 16 ≤ 333 tok/s ≤ 23.7k tok/s DGX Spark DeepSeek V4 Flash (Q2) ≤ 83 tok/s 7 ≤ 154 tok/s ≤ 13.1k tok/s Apple M5 Max 128GB Qwen3.6-35B-A3B ≤ 336 tok/s 50 ≤ 1,006 tok/s ≤ 41.0k tok/s Apple M5 Max 128GB Gemma 4 26B-A4B-it ≤ 271 tok/s 66 ≤ 1,326 tok/s ≤ 53.3k tok/s Apple M5 Max 128GB DeepSeek V4 Flash (Q2) ≤ 188 tok/s 22 ≤ 446 tok/s ≤ 29.5k tok/s Two things stand out. First, with capacity held equal, the higher M5 Max bandwidth lifts the single-session ceilings in proportion to R, and the batched ceilings by even more, because the extra bandwidth also lets more sessions clear the per-session floor. The bandwidth-rich machine wins exactly where the theory says it should, in batched throughput. Second, the memory-power column sits one to two orders of magnitude above the KV-aware column. That gap is the cost of private context traffic and expert diversity, and showing it is the point of the derivation.
Forced concurrency
What if concurrency is fixed by policy rather than chosen at the floor-satisfying optimum? On DGX Spark, pushing past b⋆ buys aggregate throughput at the cost of per-session rate.
Model Batch b Aggregate ceiling Per-session ceiling Qwen3.6-35B-A3B 32 ≤ 394 tok/s ≤ 12.3 tok/s/session Qwen3.6-35B-A3B 64 ≤ 478 tok/s ≤ 7.5 tok/s/session Gemma 4 26B-A4B-it 32 ≤ 430 tok/s ≤ 13.4 tok/s/session Gemma 4 26B-A4B-it 64 ≤ 575 tok/s ≤ 9.0 tok/s/session This is the serving tradeoff in numbers. For DGX Spark under these assumptions, 32 and 64 sessions are too high if the goal is around 20 tok/s/session, exactly the regime the usable-batch correction is built to reject, and the reason b⋆ for these models settles near 16.
A sanity check against a real run
A reported DGX Spark run served Gemma at concurrency 16 at roughly 16 to 18 tok/s/session, an aggregate of 16×16=256 to 16×18=288 tok/s. The KV-aware aggregate ceiling for Gemma at this batch is ≤333 tok/s, so observed throughput is
333256≈77%to333288≈86%of the simplified ceiling. That is close enough to suggest the implementation is near the memory-side roofline. It does not prove the quantization is optimal. The bound omits compute, scheduler behavior, kernel details, and exact cache traffic, and proving optimality would require profiler evidence of bandwidth saturation with no compute, scheduler, or CPU stalls. A bound this close simply means there is little memory headroom left to capture.
Omitted rooflines
The memory roofline is one limit among several. Real throughput is the minimum over all of them,
Tmax≤min(Tmemory, Tcompute, Tkernel, Tscheduler, Tinterconnect),and the memory term we computed can be undercut by compute throughput and tensor-core utilization, dequantization kernels, attention kernels and KV layout, prefill/decode phase mixing, scheduler overhead and request churn, CPU and PCIe involvement, multi-GPU communication, allocator fragmentation, thermal and power limits, tokenization and sampling, speculative rejection rates, and prefix-cache hit rates. Each belongs as its own limit term. The memory model remains useful because it makes the first unavoidable ceiling explicit and cheap to compute, and because for memory-bound decode it is usually the binding one.
One recurring caution applies to models with recurrent or linear state. A model with tiny fixed state and tiny private read traffic produces an enormous memory-side aggregate at high concurrency, because almost nothing in the denominator grows with the batch. That is precisely the signal that compute, kernel, scheduler, and recurrent-state details must be added before the aggregate number is treated as realistic. The memory bound describes the best case the hardware allows, and reaching it is the implementation’s job.
Cheat sheet
This section collects the whole formulation in one place, so it can be read on its own. A machine is three numbers and a model adapter is five, with the workload adding three assumptions.
Symbol Meaning C, R, O Usable memory capacity, sustained memory bandwidth, runtime overhead Wresident Full resident weight footprint, which must fit in memory Wbatch(b) Shared weight traffic per iteration, equal to Wresident for dense models and growing with b for MoE Kalloc(Lalloc) KV memory reserved per session Kread(Lread) Private context traffic per output token ρ Tokens emitted per session per iteration, one for ordinary decoding Lalloc, Lread Reserved and average read context, Lread≤Lalloc r⋆ Minimum useful tokens/s per session Everything descends from the memory roofline. If each output token must move at least q bytes and the machine delivers at most R bytes per second, then
Tmax≤qR.The first gate is whether sessions fit. Load the weights, reserve the overhead, and divide what is left by the per-session KV allocation.
bmem(Lalloc)=⌊Kalloc(Lalloc)C−Wresident−O⌋.Per-token traffic is the shared weight sweep amortized over the batch plus the private context read, which no batch size amortizes.
q(b,Lread)=bρWbatch(b)+Kread(Lread).The roofline gives the aggregate ceiling at batch b, and dividing by the batch gives the per-session rate. The aggregate rises with b while the per-session rate falls.
T(b)≤q(b,Lread)R,r(b)=bT(b)=Wbatch(b)+bρKread(Lread)ρR.Imposing the floor r⋆ rejects the batches where every session crawls, and the usable batch is whichever gate binds first.
brate(Lread,r⋆)=⌊ρKread(Lread)ρR/r⋆−Wactive⌋,busable=min(bmem, brate).The usable batch set holds the batches that fit and clear the floor, and the KV-aware bound is the best aggregate over it.
B={b:1≤b≤bmem, bq(b,Lread)R≥r⋆},Tmax≤b∈Bmaxq(b,Lread)R.Setting b=1 in the same formula gives the single-session ceiling. Dropping the private context term and taking the largest fitting batch gives the looser memory-power ceiling,
Tmax≤ρKalloc(Lalloc)WactiveD(1−CWresident+O),D=CR,and the three levels always order the same way.
actual throughput≤KV-aware bound≤memory-power bound.Every number these formulas produce is an upper bound built from memory capacity and bandwidth alone. Real implementations land below it, and compute, software overhead, and interconnects can only lower the ceiling further.
When reporting a result, always state the assumptions that move it, namely Lalloc, Lread, ρ, r⋆, the weight precision, and the KV-cache precision or attention adapter. Without them, a single tok/s number is not reproducible.
To see these bounds computed live for hundreds of audited model profiles against a catalog of local hardware, try the Local Frontier calculator.
-
I'm tired of AI writing 👃 smell 👃 here on X (and also the docs I generate) so I made a skill to 👃 de-smell 👃 Steal the skill, especially those who use GPT 5 to generate their long-form posts, for example 🤗 @analogalok @sudoingX @aijoey @ishaansehgal @stevibe and even @TheAhmadOsman 🤗 Surprisingly, @levelsio, a.k.a. AI reply guys' final boss, is the most *human* poster in the "sentence flow" metric that fable came up in an autoresearch loop. Very cool (though I haven't sampled everyone, just long-formish writers my xtap extension has picked up) I know I'm doing a favor to slop vendors out there. But there are people who are actually doing legit work, but then passing all their thoughts through GPT before putting it out here. I *have* to follow them due to my occupation. If this will save me from another "it's not X, it's Y", or "it is A, B and C", I will do it 😤 My de-smelled slop post has all the info, graphs and all: The skill (will keep updating it): This is not the end of this work, I am just beginningImage hidden -
don't tease us 😩 -
I've started using Cursor since a few days now, and I have to say the experience is really pleasant. I had not touched it since Claude Code came out in May 2025, more than 1 year ago. I had even deleted it fully last year as an editor and replaced it with Zed, after it had become very bloated and unstable I can say it no longer feels bloated/unstable and that they have transitioned successfully to the agentic era Using it mostly for fable. Using it locally in the desktop app, and also through acpx, to make codex ask fable for review Thank you @cursor_ai for sponsoring my Ultra plan! -
I told it only once before to do it on the session, and now codex habitually asks fable for review on data models and plans it creates through acpx (fable inside cursor) I then asked it what it thinks about the feedback from fable. gpt-5.6-sol is not impressedImage hiddenImage hidden -
Reminder that this fever dream of a podcast exists with @dwarkesh_sp and @_sholtodouglas youtube.com/watch?v=3XDad4…Image hidden -
-
Building an AI de-smeller
Note: this post is fully AI generated, written by Claude Fable 5 through Cursor. It was written interactively, in a back and forth with an agent, under the very kill-ai-smell skill it describes, as a demonstration of that skill.
Everyone knows the feeling by now. You open a page, read two paragraphs, and something tells you a model wrote it. The tells have become cultural shorthand, with the em dash as the poster child, but most of the discourse stays at the level of vibes. I wanted to know whether the feeling corresponds to anything you can measure, so my agent and I ran a small stylometric study, and the answer turned out to be a clear yes. A handful of surface metrics, all computable with regular expressions and a sentence splitter, separate generated copy from human writing by an order of magnitude, in several cases with no overlap between the groups at all.
This post walks through the corpus, the metrics, and the numbers, and ends with the kill-ai-smell skill that came out of the exercise.
The corpus
For the AI side I needed pages that read as generated in the wild, and I had convenient specimens close to home. Ten project sites from the OpenClaw ecosystem (crabbox.sh, mcporter.sh, gitcrawl.sh, clawpatch.ai, fs-safe.io, spogo.sh, imsg.sh, wacli.sh, gogcli.sh and goplaces.sh) have landing copy written largely by agents, and I decided to keep those pages as they are and use them as data. Their prose comes to 4,853 words after stripping code blocks. These pages were written by GPT 5.5, so the measurements characterize that model’s copy. They do not necessarily generalize to other models, whether Claude or open-weight ones like GLM and Kimi.
For the human side I wanted texts that are provably human, which means they had to be frozen before language models could have touched them. Five are essays and documentation: the SQLite testing documentation, Joel Spolsky’s 2000 essay “Things You Should Never Do”, a 2018 antirez blog post, Paul Graham’s 2009 “Maker’s Schedule, Manager’s Schedule”, and Julia Evans’ 2019 “Get your work recognized: write a brag document”. The other three attack the register objection directly, because landing copy and essays are different animals regardless of author. They are the ripgrep README at its 2016 tag, the Redis README at 3.2.0 from the same year, and the Requests README at v2.13.0 from 2017, each taken at an old git tag whose commit history proves its date. A README that sells a tool is the fairest comparison for a landing page that sells a tool. The human set comes to 15,317 words, and every rate below is normalized per 1,000 words so the corpora are comparable.
The exact texts I measured are archived in the ai-smell repository and listed in the appendix, so anyone can rerun the numbers against the same input.
The metrics
The measuring script strips code, splits sentences, and counts things. There is no model in the loop and no judgment call in any metric. That is the point of the exercise. If the tells are real, they should be detectable by grep, and anyone should be able to reproduce the numbers. The scripts, the corpus, and the figures live in the ai-smell repository.
The chart below shows every document against every metric, with the AI pages in orange and the human baselines in blue. The prose after it sticks to the ratios, because the ratios are the story, and the raw ranges are collapsed underneath for anyone who wants to check them.
Raw ranges per metric
Metric AI set (10 docs) Human set (8 docs) Em dashes /1k words 0.0–61.3 0.0–4.7 Exactly-three lists /1k words 6.3–15.9 0.0–2.0 Labeled bullets, % of all bullets 53–100% 0–11% Fragment sentences (≤4 words) 3.6–41.9% 1.4–17.4% First person /1k words 0.0–2.1 0.0–50.9 Type-token ratio (first 280 words) 0.59–0.69 0.53–0.67 MTLD lexical diversity 112–229 66–146 Mean Zipf word frequency 4.86–5.30 5.28–5.99 Sentence flow (mean run percentile) 0.19–0.41 0.49–0.73 The em-dash gap is the famous one, and it is real but weaker than its reputation. Averaged over each corpus, the AI pages use em dashes at roughly eighteen times the human rate, and the heaviest page lands one every sixteen words. But the tell only works in one direction. Three of the ten AI pages use fewer em dashes than the 2016 ripgrep README, and one uses none at all. A page drowning in dashes is almost certainly generated. A page without them proves nothing.
Exactly-three lists (“A, B, and C”) turn out to be the better punctuation-level tell. Every AI page produces them at least three times the rate of every human text, with no overlap anywhere. Even the most triad-prone human text, Joel’s essay, sits at a third of the most restrained AI page. Averaged over the corpora the gap is about nineteen-fold, and unlike the em dash, no AI page escapes it.
The labeled bullet was the discovery of the study for me. It is the bullet that opens with a short label, then a period, colon, or dash, then one sentence of elaboration. Here is a real one from mcporter.sh:
Typed clients. mcporter emit-ts emits
.d.tsinterfaces or a ready-to-run client wrappingcreateServerProxy()so agents call MCP tools with full TypeScript types.The metric is the share of a document’s bullets that follow this shape. On the AI pages it is roughly four of every five, and on one page every single bullet does it. In the human baselines the share never reaches one in eight, and five of the eight human texts never use the shape at all; their bullets are plain items, like file names or flags, without the label-and-elaboration mold. If you have seen an AI-written landing page, you have seen walls of these, and it turns out the wall is more diagnostic than the punctuation.
Fragment sentences, the verbless punches of four words or fewer, sit in between. Most AI pages run high, and the worst writes two of every five sentences that way, but the groups overlap. The deliberately punchy Requests README out-fragments three of the AI pages. Fragments corroborate rather than convict.
First person taught me the opposite lesson. It looked like a strong tell until the corpus got fairer. Against essays the gap is enormous, since the antirez post averages more than one first-person word per sentence and the ten AI pages together contain exactly one. But the pre-LLM READMEs behave like the AI pages here. The Requests README has no first person at all, and ripgrep and Redis barely any. Authorial voice turns out to track register rather than authorship, so it works as a confirming signal at best. Vocabulary variety weakened the same way. Generated copy rotates in a fresh synonym at every mention, which pushes its lexical diversity high, and the AI pages do cluster at the top of the range. But the Requests README, which is deliberately punchy marketing prose, scores right among them, so the metric separates registers more than it separates authors. The human tendency it gestures at is still real, though. Human writers reuse the established word for a thing and repeat phrases for emphasis. Joel opens three consecutive paragraphs with “You are throwing away”, and a model would never.
The vocabulary story does not end there, because the type-token ratio was the wrong instrument. We remeasured word choice with two better ones. MTLD scores lexical diversity in a way that does not depend on document length, and the wordfreq package places every word on the Zipf frequency scale, where “the” scores about 7.7 and anything under 3.0 sits outside roughly the 30,000 commonest English words.
Both separate the groups where the TTR could not. All ten AI pages score above 111 on MTLD while seven of the eight human texts stay under 106, with the Requests README as the lone crossover once again. Mean word frequency is nearly as sharp from the other side. Every AI page averages Zipf 5.30 or below and every human text 5.28 or above, so the two ranges overlap only in a sliver 0.02 wide, where the ripgrep README brushes past the plainest-worded AI page. Neither axis classifies alone, since ripgrep crosses the frequency boundary and Requests crosses the diversity one, but each README fails only one test. Draw both thresholds, at Zipf 5.35 and MTLD 100, and every AI page sits in the rare-and-rotating corner while no human text enters it. The structural detector coming up needs only one of its two lines; this pair needs both, and both suffice. The reason is not exotic vocabulary. The rare words on both sides are ordinary jargon, “OAuth” and “stdout” against “Valgrind” and “malloc”. What differs is the connective tissue. Nearly half the words in the human texts are the commonest ones in English, the “the” and “of” that full sentences run on, while on the AI pages that share drops below three in ten, because telegraphic noun piles need no articles. So the rarity metric measures, from a third angle, the same thing the fragments and the labeled bullets measure, which is that generated landing copy does not write whole sentences.
Flesch-Kincaid grade, the standard readability score, misses all of it. The AI pages come out as easier reading than the human baselines because their sentences are short, and the formula never looks at what fills them. This post sits on the human side of both word-choice axes, as the green diamond in the chart shows.
Structural tells
Beyond the counters, we tested two structural claims, and both held on the larger corpus.
The first is identity deferral, which I wrote about in Good READMEs say what tools are. Generated copy describes what a tool does and dodges saying what it is. In sentences where the tool name is the grammatical subject, action claims outnumber identity claims five to one across the ten pages. The raw ratio alone proves little, since prose about a known subject is naturally verb-led. The positional version is the tell. All three pre-LLM READMEs establish identity in their opening lines: “ripgrep is a line oriented search tool”, “Requests is the only Non-GMO HTTP library for Python”, and Redis opens its first section with the literal heading “What is Redis?”. Among the ten AI pages, exactly one does the same. Four open with headless fragments like “A local-first GitHub triage tool for maintainers and agents.”, which name a category but carry no subject and no verb, and the remaining five open with benefit imperatives or setup instructions like “Keep your editor and git workflow.”
The second is heading register, and it produced a fun inversion. Title Case, the thing style guides nag about, belongs to the humans here. The SQLite docs use it in a fifth of their headings, as was the convention of their era, while the AI pages write modern sentence case throughout. What convicts the AI headings is rhetoric. About a third of the AI headings are slogans, imperatives, or rhetorical frames rather than labels; in the human set the share is one in ten, and those are mostly mild FAQ-style questions like “Why should I use ripgrep?”. The strongest single shape is what I now call the comma couplet, a parallel two-beat slogan like “Local loop, remote box”, “Two jobs, one binary” or “Small surface, clear split”. It appears eleven times across five of the ten sites. The human set produces it exactly once, and the exception is instructive. It is the title “Maker’s Schedule, Manager’s Schedule”, a deliberate one-off rather than a house pattern stamped down a page.
There is also a tell that only becomes visible when you put the pages side by side. Six of the ten sites have a “Pick your path” section, seven make a “five minutes” time-to-value promise, eight have a “Status” section, and six close with the exact sentence “Released under the MIT license.” Each page looks fine alone. Together they reveal one prompt’s house style stamped across unrelated projects.
A minimal detector
The expanded corpus simplified the detector, because the two structural metrics turned out to need no help. Flag a page as AI-flavored when exactly-three lists exceed 3 per 1,000 words or the labeled-bullet share exceeds 30%. Either rule alone classifies all eighteen documents correctly. The em dash dropped out of the detector. It convicts a page when present in bulk, but three of the ten AI pages use fewer em dashes than the 2016 ripgrep README, so its absence clears nothing.
Plotting the two structural metrics against each other shows how much margin the thresholds have. Every human text sits in the bottom-left corner, below both lines, and every AI page sits far outside both.
Eighteen documents make a demonstration rather than a validated classifier. The first version of this study used only essays and documentation as baselines, which left the objection that the metrics were separating registers rather than authors, so the corpus now includes three pre-LLM READMEs that sell tools the way the AI pages do. The structural gaps survived that control untouched, while two metrics that looked strong against essays alone, first person and lexical diversity, collapsed into register signals. That is the argument for keeping the baselines adversarial. The next escalation would be a large sample of post-LLM, human-written landing pages, but the sizes of the surviving gaps, three-fold at the closest edge and roughly twenty-fold on average with no overlap, make me confident the structural metrics would hold.
The flow of a sentence
Everything above counts features one at a time. The last metric came out of a different question. Read the two corpora side by side and the sentences feel different in a way none of the counters capture, so I asked whether the rhythm of consecutive sentences could be measured too. I did not know what to count in advance, so we searched for it in the style of karpathy/autoresearch. A frozen harness feeds every document to a candidate scoring function as a bare sequence of per-sentence measurements and reports how cleanly the scores split the ten AI pages from the eight human baselines. The scoring function is the only file that changes between runs, and every attempt lands in a journal that now holds over fifty experiments, most of them failures.
The failures narrowed the answer. Statistics of pure order, which measure how sentence lengths rise and fall while ignoring how large they are, all fell short of separating the groups, so the tell is not in the alternation. What survived is almost embarrassingly simple. Split each sentence at its punctuation marks and keep the longest piece, the longest run of words the sentence lets through without a pause. Human writers keep producing sentences that contain one long run, whatever their register. The AI pages break nearly every sentence before a run can develop.
The first version that separated the corpus scored each sentence on a ramp between a 10-word run and a 15-word run, and it worked, but both constants were picked by hand. Ranking removes them. Give each sentence the fraction of all runs in the corpus that are shorter than its own, and average those percentiles over the document:
flow=n1i=1∑nF(ri)where ri is the longest run in sentence i and F is the distribution of runs pooled over the whole corpus. Statisticians will recognize the Mann-Whitney rank statistic. The flow score is the probability that a random sentence-run from the document outlasts a random run from the corpus, and it carries no tuned constants at all. We also tried a sliding-window generalization that multiplies the percentiles of consecutive sentences, and it looked stronger until we normalized the scale, at which point the gain vanished. That correction is in the journal too. The order of the sentences adds nothing; the length of the runs carries the whole signal.
The raw material makes the tell visible before any formula does. The chart shows the longest run in each of the first sixty sentences of one human baseline and of crabbox.sh, one of the two AI pages nearest the human range. The human post keeps clearing ten words without a pause. The AI page almost never does.
Averaging the percentiles gives one number per document, and the groups separate completely. The AI pages score between 0.19 and 0.41, the human texts between 0.49 and 0.73, and refitting the threshold with any single document held out still classifies all eighteen. The margin is thinner than the detector’s, since the strongest AI page comes within 17 percent of the flattest human baseline, the ripgrep README. This post scores 0.56, in the middle of the human range.
The reference implementation is analyze_flow.py in the study repository, and the search that produced it, including every dead end, is preserved in the autoresearch directory.
Long-form tweets in the wild
The corpora above have ground truth, which is what makes the thresholds checkable. The place people actually want a detector is the feed, where there is none, so as a last exercise we pointed the same counters at tweets. I keep a personal archive of tweets captured while browsing, about 27,000 at the time of writing, and from it we built one sample per account, made of every original long-form tweet (over 280 characters) in date order, for every account with at least 2,000 words of such text. That produced 42 accounts, my own included, and each sample is archived in the ai-smell repository with a source link per tweet.
The chart puts the 42 samples over three metric pairs, with the ground-truth pages and baselines left in faintly in every panel so the wild samples can be read against the corpus that set the thresholds. The first panel repeats the detector chart exactly, same axes, same limits, same two thresholds. Every panel also carries one more point, a green diamond for this post itself, measured the same way as everything else.
Nothing here is a verdict, since none of these samples has a known author process. Under the unchanged rules of the first panel, seven of the 42 accounts trip the detector. Four cross the triad line, three cross the labeled-bullet line, and none cross both. On the landing pages the tells came bundled, every page far past both thresholds at once, while in the feed each flagged account trips exactly one rule. The bullet share also rests on smaller counts here, since a thread carries far fewer bullets than a landing page, and one of the three flagged accounts crosses the line on just two labeled bullets. So the thresholds transfer, but the confidence does not; a verdict in the feed would need feed-specific baselines.
The other two panels show where the feed does and does not resemble the corpus. On the dash axes, six accounts run past every human baseline, topped by one at 34.5 dashes per 1,000 words, denser than eight of the ten AI pages; but the dash already dropped out of the detector on the ground-truth corpus, and it stays out here. Fragment share spreads the accounts across the whole range the two corpora span, so it separates nothing in the feed either.
The flow metric from the previous section gives an independent read on the same samples. It puts 13 of the 42 accounts below its midpoint threshold, and all three accounts that cross the labeled-bullet line are among them. That agreement is worth pausing on, because the two measurements share nothing. One counts bullet shapes and the other reads clause lengths, yet they flag the same accounts. The two lowest scorers sit below every AI landing page in the corpus, and the usual register caveat applies with extra force here, since the threshold was calibrated on landing pages against long-form prose and a punchy feed style will read as low flow on its own.
The account whose long posts sent me down this path is one of the three past the bullet line. Its sample writes 41 labeled bullets out of 90, a 46% share, half again past a threshold that no pre-LLM baseline came near, while its dash and triad rates sit mid-field. The counters flag it rather than clear it, with the caveat above about register. Its long posts also share a hook template the counters never measure. They open with lines like “90% of AI developers just…”, pivot on “It’s not X. It’s Y.”, and close with “Let me explain.” That is the next metric worth building, and the tweet samples are archived so anyone can beat me to it.
The de-smeller
The practical output of all this is kill-ai-smell, a skill in my tools repo that any coding agent can load. It covers the tells from punctuation up through page structure, and every rule carries a bad example next to a rewrite, because the rewrite teaches the hard half of the lesson. Fixing a smell means restructuring the sentence. Swapping an em dash for a punchy colon changes nothing.
The most instructive moment of the project happened while writing it up. My agent, fresh off measuring contrast rhetoric as a top tell, produced a report whose highlighted callout was titled “The strongest tells are structural, not lexical”. The detector fires on its own author. That lesson is now the second paragraph of the skill. Knowing the rules is no defense, because these patterns are how models write by default, so the sweep has to be mechanical, applied to your own output, and repeated on the text that describes the sweep.
Feel free to steal the skill, and if you run the metrics on a corpus of your own, I would love to see the numbers.
A plea
I realize I am doing slop vendors a favor by putting this out there. The tells are enumerated and scripted now, and anyone who wants their generated copy to pass can point a model at this post and patch the fingerprints. I accept that trade. What I can’t bear is actually legit people using AI to write their long-forms and publishing the model’s house style untouched, and I’d rather have everyone, slop vendors included, stop using patterns like “It’s not X. It’s Y.” and the triads than have to read them ever again.
I also know there is a brain behind those AI-generated posts. Somebody lived the experience worth posting about, formed the opinion, and then let a model flatten it into the same shapes as every other post in the feed. This favor is for them. The de-smeller is there so they can keep using the machine without shipping its defaults.
Appendix: the corpus
Every text is archived as measured in the ai-smell repository, which is the maintained home of the study. It holds the corpus, the analysis scripts, the raw results, and the figures. The AI pages were captured from the live sites in July 2026, with code blocks still in place (the script strips them before counting).
The corpus splits into four groups there:
- corpus/ai holds the ten OpenClaw landing pages.
- corpus/human holds the eight pre-LLM baselines: the SQLite testing docs, essays by Joel Spolsky (2000), antirez (2018), Paul Graham (2009), and Julia Evans (2019), and the ripgrep, Redis, and Requests READMEs at their 2016–2017 git tags.
- corpus/tweets holds the 42 long-form tweet samples, one file per account, date-sorted, with a link back to each tweet.
- corpus/self holds this post itself, archived as measured, provably AI-written by its own disclaimer. Written under the kill-ai-smell skill, it clears the detector from the other side, with zero em dashes, exactly-three lists at a sixth of the threshold, and no labeled bullets. The tells are a default, not a fingerprint, and a model instructed against them stops producing them.
-
-
Anyone else notice that any non-codex harness is more token-efficient than codex on the same task? -
I SEE BURNS... BURNS EVERYWHERE -
Here is the conversation where I built it. You can just prompt things @ratatui_rs highly recommended for TUIs. It just looks nice github.com/dutifuldev/ann… -
I vibed a TUI in just 13 codex messages, and it works great -
-
-
.@WisprFlow is much more convenient on android than iOS because the accessibility settings are more permissive It lets me overlay a button to record, anywhere on the screen, and lets me keep the regular keyboard at the same time -
-
this is cool agentic behavior under /goal. in a separate session, I had created an implementation plan in the main branch. The agent picked it up without me explicitly telling it to. The goal was to "$autoimplement finish completely" Broad orders running in loops are usefulImage hidden -
huge endorsement by geohot on GLM 5.2 -
This is significant, not sure how it will play out. There is a chance it might be for the better -
More on the way!!! -
yo @huggingface paris hq has a SICK gym to celebrate, I invented some 🤗 exercises for you mandatory for every employee -
herdr is going places -
Experimenting with a credential broker, and a telegram based approval flow for giving your claw access to your main @huggingface accountImage hidden -
1.5 TB UNIFIED MEMORY MAC STUDIO PLEASE GOD -
OKF looks interesting -
Worried that giving your 🦞 @openclaw agent write access to your 🤗 @huggingface account can risk deletion of your datasets/models/spaces/buckets, or cause irreversible damage? 😱😱😱 No need to be! I have created an agent login helper to run in your YOLO mode remote machine, which prevents any risk of irreversible deletion. Just run: uvx hf-auth-helper agent login There is a specific set of scopes you can choose while creating a fine-grained HF token. These include all read scopes + discussion.write, which let's your agent create PRs. Since you are not giving repo.write, your agent cannot force-push your main branch, change repo settings or delete them It is unfortunately not super straightforward to choose those on the web UI. This will hopefully change soon, and this functionality might even be natively in hf cli Until that happens, use hf-auth-helper to login worry free in your remote or local openclaw instance Your agents will be able to create PRs on datasets/models/spaces, which you will then be able to merge on your own browser 2 caveats: 1) This does not solve the data exfiltration attack vector—nothing does. Make sure to exclude any repos which absolutely must remain private while choosing your scopes. See for more info: 2) As buckets are not repos, your agent will not be able to modify a bucket (add/remove data). To help with that, I have another project on the way, a credential broker. Stay tuned, coming soon Source: Demo authentication flow: -
-
OpenAI is following the ASPAVA strategy, with gamified subsidies ASPAVA is a type of kebab shop from turkey. These shops usually serve mid quality food But despite that, the majority loves them. That's because they serve free stuff throughout They serve free sides at the beginning. They serve free sides during the main course. They serve free dessert at the end. They serve free tea as many times as you like Some of them give you so much "free" food that you feel like you are stealing. A lot of people frequent these shops not because they like it a lot, but because they are addicted to getting things for free Main dish portions are made smaller in order to breakeven with so many side dishes, which I feel is analogous to what is happening with GPT these days---though I can't prove it NVIDIA is not a car, and OpenAI is not kebab. I don't know if codex resets cost OpenAI much, but if they are, then they might be a waste of resources Developers are the most disloyal customer group. Once the subsidies are gone, they can switch away in the blink of an eye -
-
-
-
Most people, including professionals, will likely NOT pay more than $10k capex for their home AI workstation, even in the long run (ignoring future inflation) If you stretch it, <= $20k They will also not want to pay more than $200~300/mo opex, on electricity for example Builders who are spending a lot more than that on hardware: keep that in mind, if you are building for the general public dogfood your product with that which everyone will have, not just 0.1% -
-
Is anyone able to run nvidia/Qwen3.6-35B-A3B-NVFP4 with the config suggested in the readme? It OOMs before it can start serving huggingface.co/nvidia/Qwen3.6…Image hidden -
I know that my macbook local model benchmarks have started when my lap catches on fire Add to this list: "does not set the room on fire" -
-
I quite like this token speed simulator by @mikeveerman And I keep losing it, so hopefully I will remember to come back to this tweet:) Link: mikeveerman.github.io/tokenspeed -
I keep seeing insanely expensive builds giving insanely impressive results These results don't matter What matters is, whether one can: - run a "SOTA level" model, whatever that is - with under 32gb VRAM or unified memory - in 5 parallel sessions - with 50~100 tok/s each - with enough leeway memory for other applications - in a system as cheap as $1000 That is our goalpost That is the threshold when open source AI will win -
-
Another annoying GPT-ism (circa June 2026): while describing something, it always describes what it *does*, but never what it *is* > "LocalPerf benchmarks local LLM inference servers and keeps the evidence in one portable run artifact." If I were a philosopher, I would say that "AI lacks ontology". Or that it is "anti-essentialist", believes that things cannot be things in themselves But all models literally have an ontology. They have it since word2vec days, you can plot it out. It's just an annoying tendency in GPT's writing So using philosophy to understand AI might be dumb sometimes If you don't want your README's to sound like slop, then you can steal my write-readme skill: After write-readme: "LocalPerf is a local LLM inference benchmark CLI. It runs benchmark plans against local inference servers and stores the evidence in one portable run artifact." Skill: github.com/osolmaz/tools/blob/main/agents/skill…Image hidden -
Give @LottoLabs a follow if you are not already He is building localmaxxing.com, crowdsourced LLM benchmark results and performance profilings Very very coolImage hidden -
This. Especially when the whole machine hangs instead of OOMing when my agent accidentally loads too many models into memory (lmk if there is a firmware update or sth that fixes this on the spark now, creating a cgroup doesn’t work) forums.developer.nvidia.com/t/dgx-spark-be… -
Apple hiked prices as I was posting this 💀 -
Was great to be there, thanks for the invite @lionelsimai! -
How I imagine @TheAhmadOsman I actually bought my GB10 thanks to him back in feb, at a discount Give him a follow if you are not already!Image hidden -
The Local Frontier is advancing The amount of AI memory for inference we can get for less than $3000 has been steadily increasing The memory crunch has slowed this down and even made it retrograde. However, once we bounce back from it, the progress will be gloriousImage hidden -
Was great to talk, thank you @ben_burtenshaw for inviting me! -
I knew it would find me, sooner or later 🫠 -
Besides being hot, this take is very correct and points out to a fundamental tradeoff in the storage layer being centralized versus distributed git is distributed and that makes total sense for code which takes small space by its nature. it is cheap for everyone to duplicate it locally. this proved to be very useful over e.g svn, when devs could develop independently from the centralized server AI artifacts, however, are 1 million times bigger than code. in that case, the bottleneck becomes storage and network. decentralization and version control become lower priority. they can be sacrificed the tradeoff tilts towards getting the cheapest possible storage and transfer. because you will need a LOT of that I regret to announce to competitors that Hugging Face has already won this game when they acquired Xet. the tech just works, and the network effects are immense -
Awesome. A feat unimaginable before agents -
-
My posts here on X now sync automatically to my blog, giving me full ownership of my content and zero effort SEO For free, no API costs My long-form posts are automatically featured and titled on the front page of solmaz [dot] io. Filtering is done by my claw running a sync-x skill daily, which then notifies me on Discord How do I scrape the posts? ALL the posts I view (including the ones I post) are saved locally using @kubmi's xTap and then synced to a private repo using my extension on it xtap-sync My claw has access to that private repo and can run programmatic tasks like sync-x, summarize what happened that day, notify me about any topics I wantImage hiddenImage hiddenImage hidden -
Excited! -
If you are interested in running such demos, look into --demo mode in my local model swiss army knife localpi github.com/dutifuldev/loc… Thank you @googlegemma for the shoutout -
i meant to qt this one x.com/herdrdev/statu… -
Huge. Added my github TUI github.com/dutifuldev/ghz…Image hidden -
Will talk about my recent adventures running local models on the Spark, Pi and OpenClaw Click Notify Me on youtube to stay tuned -
New blog post: Using local models for agentic zero-shot classification, in real-time, high frequency triage If you have a 128gb of memory for models (a DGX spark like I do for example), you can create a real time classifier and notifier for yourself that can classify more than >20 items per minute, using mid-sized @googlegemma and @Alibaba_Qwen models, with over 200-300 output tok/s aggregate throughput Like processing new tweets on twitter, issues/prs on github, messages on telegram and discord, in real-time Over the past few weeks, I have built one for myself, to filter and get notified about local model related issues on the OpenClaw repo I initially thought gemma-4-e4b would give me the best tradeoff I was wrong. I learned that if one has enough memory already, one should not bother with <10b models like gemma4 e4b or e2b. Precision and recall were much higher zero-shot with gemma-4-26b-a4b, whereas the smaller e4b needed significant prompt optimization to eventually not perform nearly as good To provide more context to the model, I created a restricted bash-like shell, called reposhell. In that shell, it can run read-only commands to ls/find/grep/cat openclaw source code, but only that. When the PR description/diffs are not clear enough as to categorize it, the agent reads the code to figure it out Because small models can get prompt injected, and I need to make sure that someone can't harm my setup by creating a malicious issue or PR in the openclaw repo I found that for specific systems like this, it is very convenient to extend and bundle Pi. You can create agentic CLI tools that work fully locally and for free, and keep that separate from your main pi coding setup. localpager-agent has its own session dir and tools, and I ensure that it will run local models in a secure way by isolating it from my main pi setup Once localpager-agent categorizes a PR/issue as local_models and related labels, I automatically receive it as a notification on Discord The whole implementation is fully open source and MIT licensed, alongside the dataset we used to benchmark the performance I believe zero-shot agentic classification running on local hardware will find many use cases across a wide variety of business applications, like news gathering, open source software development, customer support, content moderation, sales and so on Agents increase the amount of information produced in a lot of systems, and hence we will need to set up cheap ways to wrangle all that information In times where governments can cut off access to SOTA models on a whim, it is more important than ever to build your business on open models and if possible, run them on your own hardware! Big thanks to @evalstate and @ben_burtenshaw for their valuable feedback, especially with helping me evaluate this more rigorously! One take-away is that categorizing contributions in an open source repo is a *hard* problem, and that it is not trivial to reliably create a golden dataset with LLMs, for evaluation purposes Read more here: huggingface.co/blog/local-models-pr-triage -
One sweep over 100 samples takes around 4 hours. Next up: cross reference ground truth with predictions from hf-mem by @alvarobartt github.com/alvarobartt/hf…Image hidden -
gpt5.5 and most other models are very bad at one-shotting nice data models gpt5.5 also has this annoying property that once it decides for a schema (or any design), it's very hard to trigger thinking again. and if you ask to "rewrite from scratch", it will write create something even more ridiculous To solve this problem, I have built a meta-harness over codex just for simplifying slop data models called schemator (work in progress) Basic idea: it mimics what I myself do while I am designing a schema: scrutinize and question each field one by one It starts a fresh codex session for each field with a fixed prompt like "Try to come up with the most Lindy data model" + a prompt for side notes It does that with a fresh context for each field, so that they are independent from each other. At the end of a review run over a field, the reviewer can propose to keep, rename or remove the field When all fields are reviewed once, that makes one iteration. Then this is looped over until the review results stabilize, and do not propose any further changes I get better results by just asking my agent to "use schemator on this" after it creates a JSON schema or SQL table Give it a try if you have codex! It has a skill, so should be easy for an agent to figure out how to use github.com/dutifuldev/schematorImage hiddenImage hidden -
gpt 5.5 is not naturally good at modeling and cannot create simplified nice mathematical models completely autonomously I did a parameter sweep with gemma-4-31b-a4b on memory usage, output tok/s etc. while varying context window, concurrency and other parameters. It took quite a few tries, and I still do not trust the model that gpt5 fit to the data besides, it measured linux cgroup memory and not the actual gpu memory used, so the whole sweep is wasted... output tok/s looks more accurate though, soon I will have a model that can give the optimal parameters over the space of context window <> concurrency <> tok/s <> memory usage off to do another run -
For my recent LLM leaderboard osolmaz-leaderboard.hf.space, I sum up all time total downloads (or likes) across model variants, and then divide it by the age of that model. I.e. "time decay" for popularity This gives a more time-agnostic metric for the popularity of that model. In an ideal ranking, older models that are not popular anymore should be demoted, like 2 year old Llama 3 models. If you don't do that, they might still occupy top 10 needlessly, despite having been replaced by e.g. qwen in practice Thanks to that, qwen-3-6b which came up 1 year ago and has 150m downloads can surpass llama-3-1-8b which came up 2 years ago and has 200m downloads More notes on my post: solmaz.io/popularity-rankingImage hidden -
if you take the Most Downloaded Models of All Time, Llama 3.1 makes it to Top 10 with around 200 million total downloads (ranking is done w.r. to time-averaged downloads) RIP Llama, you walked so @googlegemma and @Alibaba_Qwen can run Also a reminder that if you build your branding on top of open weight models developed by big corps, you might eventually be the de facto owner of that brand if they pull the plug on it. Like llama.cpp @ggml_org Huge fumble by MetaImage hidden -
My LLM leaderboard osolmaz-leaderboard.hf.space auto discovers different variants of model releases, even if they are not linked by base_model From this, I found out that @RedHat_AI was the first to release NVFP4 quantization for qwen3-6-35b-a3b Nice to see everything in one placeImage hiddenImage hidden -
"big token will bless me with free tokens today inshallah" is not a healthy mindset to nurture don't get me wrong, I love the subsidies and the memes@thsottiaux ·Image hidden -
Popularity ranking cheatsheet
I recently did some work ranking models on Hugging Face. While doing that, I remembered some concepts I had known years ago from studying recommender systems. But I couldn’t find any personal notes from that time.
So I’m leaving this cheatsheet for my future self, if I ever need it again.
The main idea with popularity metrics is that it is proportional to e.g. total likes/views, and inversely proportional to the time passed to accumulate those likes/views. A lot of different platforms came up with many different ways to calculate this. And while you can model this in a certain way that maximizes some imaginary objective, what ends up being implemented first is the cheapest/most efficient algorithm.
Below are some examples, generated by GPT 5.5 xhigh.
<slop>The abstract problem is:
rank items by scarce attentionA platform has many items and a limited front page. It needs to decide what deserves visibility now. That is usually not the same as “best,” “most useful,” or “most popular all time.”
A clean taxonomy:
popular=received a lot of attention hot=received a lot of attention recently trending=receiving more attention than expectedHere A means attention: views, downloads, likes, votes, streams, sales, stars, comments, clicks, etc.
Raw popularity
This is the simplest ranking.
S=AUse it when you want “biggest ever.”
Examples: most downloaded, most viewed, most sold, most starred.
Problem: old items dominate because they had more time.
Velocity
This measures speed of attention.
S=tAwhere t is age.
Use it when you want “how fast is this spreading?”
A stricter version:
S=(1+t)αAIf α=1, this is close to attention per unit time. If α<1, old items are penalized more gently. If α>1, new items are favored aggressively.
This family is close to what Hacker News describes at a high level: HN says its basic ranking divides points by a power of time since submission, while also applying other factors such as flags, anti-abuse systems, demotions, account/site weighting, and moderator action.1
Log-scaled velocity
Raw attention often follows a power law: a few items get enormous numbers. So platforms often compress the signal.
S=(1+t)αlog(1+A)This keeps huge items ahead, but prevents them from crushing everything else.
This is usually a better “hotness” formula than plain:
S=tAbecause it rewards scale without making scale the only thing that matters.
Recent-window popularity
Instead of lifetime attention, count only a recent window.
S=Arwhere Ar is recent attention.
Or normalize by window size:
S=wArwhere w is the time window.
Examples:
most viewed today most streamed this week most downloaded in the last 30 daysSpotify’s daily and weekly charts are this kind of family, though Spotify also says it uses chart-eligible streams and filtering formulas to protect chart integrity; it does not simply expose raw app stream counts as chart counts.2
Momentum
Momentum compares the current period with the previous period.
S=Ap+1Ar+1where Ap is previous-period attention.
Example:
S=downloads last week+1downloads this week+1This finds things that are accelerating.
Problem: small items can look extreme. Going from 1 to 20 is a 20× jump, but it may still be tiny in absolute terms.
A safer version mixes ratio and volume:
S=log(1+Ar)Ap+1Ar+1
Trend detection
Trending is not just “popular.” It usually means “unusually active relative to expectation.”
S=E+1Ar+1where E is expected attention.
If something normally gets 100 views/day and now gets 10,000, it is trending. If something normally gets 10 million views/day and now gets 10.5 million, it is popular but not necessarily trending.
Another version:
S=Ar−EThe ratio version favors surprise. The difference version favors large absolute surges.
Google Trends is a useful example of normalization: it divides search interest by total searches for the relevant geography and time range, then scales results from 0 to 100, so large regions do not automatically dominate raw volume rankings.3
Hotness
Hotness combines attention and freshness.
A simple hotness score:
S=log(1+A)−λtPopularity pushes up. Age pulls down.
Another common form:
S=(1+t)αlog(1+A)This says: “large attention matters, but old attention decays.”
Classic Reddit-style hotness
The old open-source Reddit code had a “hot” formula based on vote balance, logarithmic scaling, and time. In simplified notation:
S=sign(u−d)log10(max(∣u−d∣,1))+45000Twhere u is upvotes, d is downvotes, and T is time since a reference epoch. This is specifically the archived open-source Reddit implementation, not a guarantee of current Reddit production ranking.4
The important idea: votes matter logarithmically, and time strongly affects ordering. This makes the ranking feel alive.
Time-decayed attention
Instead of using age directly, you can make every attention event fade over time.
S=∑Aie−λΔtiwhere each attention event Ai contributes less as it gets older.
Plain language: a view today counts more than a view last month.
This is good when you have event-level data.
A simpler approximate version:
S=Ar+βApwhere 0<β<1.
Example:
S=attention this week+0.5×attention last weekSteam’s real-time Top Sellers use this general idea in a revenue context: Steam says it rolls up player spending from the trailing 24 hours and gives extra weight to spending in the last 3 hours, across base game purchases, DLC, and in-game transactions.5
Quality-adjusted popularity
Sometimes attention alone rewards clickbait. So platforms mix attention with satisfaction.
S=Aqwhere q is a quality signal.
Examples of q:
like rate rating completion rate return rate positive vote shareYouTube Charts disclose this kind of multi-signal logic: they consider view count, how quickly views are growing, where views come from, topic, age, and performance compared with recent uploads from the same channel; YouTube explicitly says the highest-view-count video is not necessarily ranked first.6
Confidence-adjusted ranking
This prevents tiny samples from winning too easily.
Bad ranking:
S=qThis lets an item with 2 perfect ratings beat an item with 10,000 very good ratings.
A Bayesian shrinkage version:
S=n+knq+n+kkqˉwhere n is sample size, q is the item’s observed quality, qˉ is the global average, and k controls how much evidence you need before trusting the item.
Plain language: with little data, pull the score toward the average.
IMDb is an example of this family in spirit: IMDb says it publishes weighted vote averages rather than raw averages, that not all votes have the same impact, and that it does not disclose the exact method.7
Wilson score ranking
For up/down votes, a common confidence-based formula is the Wilson lower bound.
Let:
p=nuwhere u is positive votes and n is total votes.
Then:
S=1+nz2p+2nz2−−−−−−−−−−−−−−−−−−znp(1−p)+4n2z2This estimates a conservative lower bound for true positive rate.
Use it when you want “best-rated with enough evidence,” not merely “highest average rating.”
Evan Miller’s “How Not To Sort By Average Rating” popularized this for web rankings, and the archived Reddit code includes a confidence sort using the Wilson method; Stack Overflow also discussed the same family of sorting methods for comments/answers.8
Category-normalized popularity
Raw popularity is unfair across categories.
S=AˉAwhere Aˉ is average attention in that category.
Example: a niche item with 10,000 downloads may be huge in its category, while a general consumer app with 10,000 downloads may be irrelevant.
A velocity version:
S=Aˉ/tˉA/tUse this for “popular relative to peers.”
Spotify’s Local Pulse is a real-world example of relative popularity: Spotify says Local Pulse shows songs uniquely popular in a city relative to their overall popularity.2
Composite ranking
Most mature platforms do not use one pure formula. They combine signals.
S=alog(1+A)+blog(1+Ar)+cq−dlog(1+t)where a,b,c,d are weights.
Then platforms add penalties:
S=S−spam penalty−abuse penalty−duplicate penaltyProduct Hunt is explicit that its homepage leaderboard changes based on upvotes, comments, time since submission, and other factors, while withholding exact details to reduce gaming.9
Amazon’s book sales ranking is also a composite/decayed-relative system: Amazon says rankings reflect recent and historical activity, recent activity is weighted more heavily, ranks are relative to other books, and rank can change even if the item’s own activity stays constant.10
Examples out in the wild
Platform Ranking type Disclosed logic Reddit Hot, classic open-source version Hotness Vote balance, log scaling, and time term.4 Hacker News Hotness / age decay Points divided by a power of time, plus flags, anti-abuse, demotions, weighting, moderator action.1 Product Hunt Launch hotness Upvotes, comments, time since submission, and undisclosed anti-gaming factors.9 YouTube Charts Trending / hotness View count, growth speed, traffic source, topic, video age, channel-relative performance, safety filters.6 GitHub Trending Developer attention The public page exposes total stars/forks and “stars today”; GitHub does not publish a full ranking formula there.11 Google Trends Normalized search interest Search interest divided by total searches for that time/place, scaled 0–100.3 Spotify Charts Stream popularity with filtering Chart-eligible streams; local charts; Local Pulse is city popularity relative to overall popularity.2 Steam Top Sellers Revenue hotness Trailing 24h revenue, with extra weight on last 3h; includes DLC and in-game transactions.5 Amazon Best Sellers Rank Decayed relative sales/activity Recent and historical activity, recent activity weighted more heavily, rank relative to peers.10 IMDb ratings Weighted reputation Weighted vote averages, not raw averages; exact method undisclosed.7
The design choices
Every attention-ranking system chooses answers to these questions:
Choice Meaning What counts as attention? Views, downloads, stars, likes, votes, sales, comments, plays, installs. Is old attention still valuable? Use lifetime totals if yes; decay if no. Do you care about speed? Use velocity or recent-window ranking. Do you care about surprise? Use trending vs expected baseline. Do you care about quality? Mix in ratings, retention, completion, votes, reviews. Do you need confidence? Use Bayesian shrinkage or Wilson scoring. Do categories differ? Normalize within category, geography, language, genre, or cohort. Can it be gamed? Add anti-spam filters, trust weighting, duplicate penalties, anomaly detection. Is it public or personalized? Public rankings use global signals; feeds use global signals plus user relevance.
Practical formula families
For all-time popularity:
S=AFor average popularity over lifetime:
S=tAFor hotness:
S=(1+t)αlog(1+A)For recent popularity:
S=ArFor momentum:
S=Ap+1Ar+1For trending:
S=E+1Ar+1For quality-adjusted popularity:
S=AqFor confidence-adjusted quality:
S=n+knq+n+kkqˉFor relative category popularity:
S=AˉAFor a practical general-purpose front page:
S=alog(1+A)+blog(1+Ar)+cq−dlog(1+t)Then apply filters and penalties.
Best mental model
There are three core ranking concepts:
Popular: S=A“Has accumulated a lot of attention.”
Hot: S=(1+t)αlog(1+A)“Has a lot of attention for its age.”
Trending: S=E+1Ar+1“Is getting more attention than expected.”
Most real platforms are combinations of these, with normalization, confidence adjustment, and anti-gaming rules layered on top.
</slop> -
If you are in AI, just don’t be anon here I see a bunch of anon accounts posting great local model content… what’s the point of being anon? To seem cool? Most of those accounts are not doing anything illegal, so there is no point. It would add so much more legitimacy to your work if you just put your real face and not a slop or anime girl pfp it only makes sense for those who are abliterating models. otherwise, it makes you seem sus just put your real face anon -
I created an LLM leaderboard based on Hugging Face download and like counts, grouped, filtered and time-averaged. Top 5 downloads is shared by @Alibaba_Qwen and @googlegemma 👑🤝👑 Top 5 likes, on the other hand also includes @deepseek_ai V4 Pro 👑 Even @OpenAI makes it to #8 top downloads with gpt-oss-20b 👑 qwen3-6-35b-a3b is the second most CIRCULATED LLM of this year, with an average of 21 million downloads per month, since the day it was released 2 months ago 📈📈📈 Despite first place belonging to 8mo old qwen3-vl-2b-instruct, the highlight belongs to the mid-sized MoE model, which has hit a size/performance sweet spot so hard that it absolutely 💥 SHATTERED 💥 Hugging Face leaderboards in the 2 months since it has launched qwen3-6-35b-a3b is followed closely by its dense sibling 27b --- and then the mid-sized gemma 4 models 26b-a4b and 31b Note that a model's distribution is inversely proportional to its size, but not strictly! Usefulness plays a factor as well, since gemma 4 26b-a4b is being downloaded more than the smaller gemma 4 e4b I created this leaderboard because Hugging Face's all time highest downloads and likes did not give me enough information about what is really popular, neither today, nor all-time. I wanted something in between How do I calculate this ranking? - Get models that with n_downloads >= 100k - Exclude models older than 1 year - Deduplicate and group quantizations and variants of the same model based on slug prefix heuristics - For each group, sum up total downloads of all time - Sort by descending total_downloads / age = average_downloads_per_day (can also sort w.r. to likes per month) - Repeat every day to get the most up to date ranking More info and source on the leaderboard page, hosted on a Hugging Face space: osolmaz-leaderboard.hf.space This is a work in progress, please reply below if you see a model that should be there is missing, or any other mistakesImage hidden -
-
gemma-4-26b-a4b is the most CIRCULATED LLM of recent history, with an average of 126k downloads per day, since the day it was released 3 months ago Top 10 is shared by Qwen and Gemma, with DeepSeek V4 Pro coming in close 🤝 Note that a model's distribution is inversely proportional to its size, but not strictly! Usefulness plays a factor as well, since gemma 4 26b-a4b is being downloaded more than the smaller gemma 4 e4b I created this leaderboard because Hugging Face's all time highest downloads and likes did not give me enough information about what is really popular *these last few months* How do I calculate this ranking? - Get models that with n_downloads >= 100k - Exclude models older than 1 year - Sort by descending total_downloads / age = average_downloads_per_day (can also sort w.r. to likes per month) - Deduplicate quantizations etc. of the same model based on slug prefix heuristics More info and source on the leaderboard page, hosted on a Hugging Face space: osolmaz-leaderboard.hf.spaceImage hidden -
I need better UI/UX on queueing messages to agents. I want to be able to: switch the order of queued messages pause the queue edit any message that are still in the queue undo steer messages in the few seconds they are being sent I want more visual emphasis on the queue, like a Queue View I can toggle, that puts the queue at the center I want this in all the UIs and coding agents, codex CLI, desktop, moshi... especially while on the phone -
saving the world from AI cartelization for fun and profit -
herdr is my terminal now, both local and remote. highly recommend -
16x parallel Gemma-4-26B-A4B-NVFP4 runs 🤯🤯🤯 18 output tokens/s, aggregate 300 tok/s 1 DGX Spark with 128 GB unified memory Concurrency so high I had to demo it programmatically It can go up to 32 even! 🤯 But then my screen would not have been readable for you And this is not even using flashinfer yet! Please reply if you know whether support is on the way Note that this is not dumb e4b or e2b that you can run on the average laptop. This is the big Gemma MoE Model link: huggingface.co/nvidia/Gemma-4-26B-A4B-NVFP4 -
-
I did some math, and running my Nvidia GB10 workstation (Asus GX10) costs me maximum: 12~13 USD / month or 150~160 USD / year It is a little bit above half the price of ChatGPT plus subscription. For that, I get to run models that can fit in 128 GB of memory How I calculated: You can see how much power your apartment uses in Singapore in half-hourly resolution. We turned off all devices and A/C while we sleep, and got only the fridge and the GB10 remaining From that, we see it uses around 80-100 Watt while I was running an inference workload overnight. So this is like an upper bound I take it as 90 Watt. Electricity here costs 0.25 SGD / kWh 0.09 * 0.25 * 24 * 30 * (SGD/USD conversion rate) = 12~13 USD / month = 150~160 USD / year Local models are getting very good now, small ones roughly around GPT 5.x-mini level. This workstation makes all sorts of workloads possible for me that would otherwise cost a ton on the API It is also my always on workstation that works overnight. I use Codex for my work, and my workstation is always running agents. It never sleeps. I never have to worry about keeping my laptop lid open. I connect and monitor the agents anytime on my phone using mosh and herdr We have crossed a threshold. Running local models is cheaper than a big token sub for quite a few workloads already. If you are running a business, that makes a difference The localening is here -
THE LOCALENING IS HERE -
New agent benchmark alert: SkillsBench -
Link to model: huggingface.co/nvidia/Qwen3.6… -
Click open GitHub PRs and issues directly in the side pane in @herdrdev, instead of having to go to the browser. As many issues and PRs as you want, WITH TABS! Install ghzinga herdr plugin and just ctrl+click the link: github.com/dutifuldev/ghzinga Thanks @lumendriada for sneaking in the ability to capture link clicks 2 days after I requested it! God I love open source... -
Hugging Face buckets are very literally, actually, 100%, a game changer note that I never ever use that phrase -
Trying to copy wrapped URLs is a pain not only in ghostty/iterm2 but also in mobile apps like Moshi On the laptop it’s fine because I can select rectangular area and edit it, or make the window bigger On the phone, its’s impossible. Fingers too big, too much of a hassle Should a terminal emulator try to detect these? It already detects herdr. What do you think @odd_joelImage hidden -
nvidia/Qwen3.6-35B-A3B-NVFP4 running in vLLM nightly on my Nvidia GB10 is actually insane 50 tok/s, 4 concurrent generations. total 200 tok/s. ideal for spawning subagents or working in parallel its tool calling behavior is very good as well. I will be giving it test drive on an openclaw instance, and keep you posted More details on NVIDIA forum: forums.developer.nvidia.com/t/benchmark-report-… -
Current average generation speeds for local DeepSeek-V4-Flash-Q2, highest to lowest: Mac Studio M3 Ultra: 32 tok/s MacBook Pro M5 Max: 30 tok/s Apple ??? M4 Max: 25 tok/s MacBook Pro M3 Max: 24 tok/s Mac Studio M2 Ultra: 22 tok/s NVIDIA DGX Spark / GB10: 13 tok/s It seems macs' higher memory bandwidth is contributing here, though I'm not sure if GB10 performance could be improved (I do hope so, I have one!) -
We have local Deep Research Now we just need to index the whole internet to have local ChatGPT 😅 -
Btw, TTS has come such a long way, @GoogleDeepMind cooked with gemini-3.1-flash-tts I gave Codex my google credentials and it oneshotted the Gemini TTS implementation When I built this 4 years ago, Azure TTS used to be SOTA. Then @ElevenLabs came in and raised the bar super high. Now Google is going after their lunch with controllable expressiveness at scale. I cheer for both! Here is Manim Voiceover demo from 4 years ago with Gemini TTS (sound on) -
-
I major concern I have these days is, while I author code in languages I cannot manually code, are they any good? Over years, I have worked with a number of languages: C, C++, Fortran, MATLAB, JavaScript But Python was my go-to language since more than 10 years. Well that changed last summer So while I have strong opinions on how Python code, should be, conventions and all, I don't have so strong opinions on other languages. That means I am producing slop by default in Rust, Go and TypeScript To solve that problem, I created github.com/dutifuldev/slophammer Its aim is to be "the only tool and resource your agent needs, to minimize slop" It is inspired by the recent bathrobe rants of @unclebobmartin, a.k.a. the author of clean code It enforces a minimum test coverage, maximum cyclomatic complexity, mutation tests, code style across different languages But I have a major issue: How do I know that Slophammer itself isn't slop? One way is to implement and use it for Python, the language I know better, and judge what kind of changes it enforces So for this weekend experiment, I used Slophammer to refactor, improve coverage and merge new features to one of my old Python projects, Manim Voiceover github.com/ManimCommunity/manim-voiceover The result is... mixed. We now have types everywhere, which is great. But the constraints have also made it write garbage code like this one. It works fine, even though it's not elegant. The new feature also works What do you think? Does code still need to be aesthetically pleasing to the human eye? Should it still be human readable? If an agent writes slop in the forest, and there is no-one to read it, is it still slop? If anything, I should use its output in Python to reason about other languages, and add more and more constraints. The more the constraints, the less the slopImage hidden -
This is why I love this site, open collaboration! -
-
This is what I have been feeling recently as well, looking at models write code better and faster than me -
Dabbling in GEPA. Codex's /goal on GPT 5.5 high is still surprisingly reward-hacking I had set a /goal before I slept to implement a plan. It ended the loop after doing just 1 iteration It feels like the model is following the path of least resistance and slacking off. Though it could also be me putting "try to make good progress in 8 hour's time" in the prompt, can't be sure Lesson: When you are doing such a solver loop, always specify min_iter and max_iterImage hiddenImage hidden -
-
-
-
-
-
Experimenting with SOUL.md on gemma4-26b-a4b (running on @DeepInfra) Interesting that such a lightweight model can already run such a conversation in openclaw harness @GoogleDeepMind cooked hereImage hidden -
-
-
-
-
What did @karpathy see / was shown? Why did the benefactor and teacher of the whole ML ecosystem join Anthropic, a company the polar opposite of his image, on the eve of such a powerful model release It can't be purely money Did he reckon that the only way to benefit humanity was to be on the inside, or rather, to not be left outside, of whatever is brewing in there? -
-
-
It’a been a little bit over 1 year since Anthropic released their Max plans and Claude Sonnet and Opus 4, thus making Claude Code affordable and kickstarting the agentic revolution Opus 4 was a glimpse into the future. I’ve spent the entire summer swearing at it and typing ultrathink Today, Fable 5 feels like another step change I no longer need to type ultrathink. And no longer need to swear at Anthropic models. Only at their marketing team. -
Fable burned through my 5 hour quota, and then automatically fell back to usage credits without asking. Org settings I suppose It was burning through 1 usd every few seconds It burned through 66 usd before I reacted. Yeah, this is not affordable for anyone with that API pricing, without subsidy/planImage hidden -
-
Speaking of loops, I have renamed my implementation-loop skill from earlier this year to autoimplement, because it's shorter Calling skills that loop auto-x, auto-y makes them more memorable than calling them x-loop, y-loop But it also increases the number of keystrokes you have to type, before you can tab-complete them Alas, I like still this more github.com/osolmaz/tools/tree/main/agents/skill…Image hidden -
-
Ok so there is auto mode which they introduced back in March, but apparently they are not so confident in it that it's still in experimental mode and not easily findable in settings code.claude.com/docs/en/permis… -
To YOLO with Fable 5, or not to YOLO, that is the question... The last time I left, Claude models still had tendencies to rm -rf your home folder or delete stuff without asking first. Is this still a risk? And from the looks of it, Claude Code still doesn't have Codex's LLM-filtered approval gate feature. Or am I missing something? Please enlighten your fellow Claude noob 😇Image hidden -
The masculine urge to create your own agent multiplexer -
We've lost another brother @maddada to agent multiplexers 🫂 x.com/maddada/status… -
Just in time for a lot of Codex-default developers going back to Claude Code momentarily to try out Fable 5 Here is a CLAUDE.md -> AGENTS.md symlinker that should save you from the hurdles of obstinate Anthropic conventions It installs a hook that creates the CLAUDE.md symlink automatically as Claude Code traverses directories that contain AGENTS.md, automatically ignored by git No need to create CLAUDE.md with reference to AGENTS.md like Anthropic suggests. It just works github.com/dutifuldev/claude-md-symlinkerImage hidden -
-
Question to my ghostty-savvy friends I am trying to reproduce the Quake style dropdown experience I have been using since 2010 on ghostty on mac here. nothing works quite as well as iterm2 yet I tried ghostty quick terminal mode. good but it doesn't let me open multiple tabs I tried cmux because it ships ghostty anyway and is supposed to have more features. but its system-wide hotkey is not playing well with aerospace and window focus iterm2 worked perfectly. tap control double and I'm in the terminal. is there anything that replicates this UX -
-
-
ghzinga can now show multiple PRs/issues in tabs natively, no need to create a new pane in tmux/herdr also, you can tell your agent to open all the relevant issues/PRs in a side pane using it, and it should work seamlessly it's the open source maintainer's best friend. life is too short to juggle 100 tabs in chrome, why not have it right next to codex! -
just vibe-checking all these models is a full-time job -
-
new open tts model, demos are eerily good -
Here is the source, I called it ghzinga. You can click click click by default (unlike gh dash, which is still awesome in itself) For just viewing single issues/PRs github.com/dutifuldev/ghz… -
-
-
@OpenAI Extra ironic that this is tweet was AI generated x.com/DuckDuckGo/sta… -
Wait did anyone think otherwise? lol 128 GB unified memory, 20 cores, "Spark" in the name... I didn't watch the presentation. Maybe because of that I directly inferred that it's the same chip -
-
F for the fallen brother 🫡 x.com/dani_akash_/st… -
RIP 🙏 x.com/yushen686/stat… -
🙏 RIP, we’ve lost another brother to agent multiplexers. Amen 🙏 -
🙏 Daily prayer 🙏 Thank you lord for giving me the restraint to not build my own agent multiplexer 🙏 Amen 🙏 -
-
we'll have a linux laptop running ds4 flash and co. !!! -
-
-
Thank you @ashleywolf for helping me personally, I really appreciate it! The account was reinstated less than 1 hour of posting this! The whole company must be working hard to make github scale in an era of crazy demand and growth! -
I am a paying customer of github. I have a team account with 2 seats, one for me, and one for my agent. I have been paying for more than a year now I do this because I treat my agent's workstation as a lower trust machine, and do not allow merging to main in certain repos I have been working on a tool that calls github's graphql API. today, my agent's account username:dutifulbob got suspended for no reason what am I supposed to do now? put my main account on my openclaw instance? I applied to reinstate, it appears it might take weeks to enable it back??? Maybe don't pull such things on your long term paying customers @github??Image hidden -
🙏 Thank you lord for giving me the resolve and patience to not build my own agent multiplexer Amen 🙏 -
-
This looks very promising! -
-
I still have 0 progress on this, did anyone else experience this bug on codex desktop app as well? Here is a theorized repro on a fork, but I am not sure because I don't know how it happened exactly github.com/osolmaz/codex/… -
You know how every time you create a new repo on GitHub, you need to: - create a branch protection rule on main, block force pushes - enable auto-merge - make branches delete when PR gets merged - enforce linear history, disable merge commits - enable update branch button Now, you can do all that with a single command: npx github-sane-defaults@latest plan <your-org-handle> --all This is just the plan command, shows you which repos don't have branch protection rules and those settings, like: changes awesomeorg/awesomerepo Settings allow_merge_commit true -> false allow_auto_merge false -> true allow_update_branch false -> true delete_branch_on_merge false -> true Ruleset create Then you run it with apply instead of plan, and it makes those changes Basically a very simple github policy manager. Let me know if you want configurability with policy files, currently it just applied my opinionated defaults Either way, this will never be something too deep, there is terraform for that Source: github.com/dutifuldev/github-sane-defaultsImage hidden -
-
-
-
-
-
Another shoutout to Agent of Empires. It's still not exactly what I want from my personal tool, but "cmux but fully in terminal" is a good long term vision imo -
Looks very promising for debugging agent sessions across different harnesses, still has some rough edges that need to be polished -
If only codex desktop app was open source, and i could put it fully in the terminal, and i could program what panes each session can show -
-
-
Btw, there is no reason not to substitute this with a maxed out Macbook Pro Max as the workstation (which gives you 128 GB memory) and a Macbook Air as the terminal device Might be more feasible for digital nomads, since GB10 is 1.5 kg without that huge adapter, and travelling with all that might raise some eyebrows -
Clankers are NOT Humans Clankers are NOT Individuals Clankers are NOT Persons NOT a Human: This is straightforward. Is it of the homo sapiens species? No → Then it is not a human --- NOT an Individual: Does the clanker have its own boundary? Does it govern itself inside that boundary? Can it defend that boundary? No, no, and no → An LLM is a file copied en masse to data center hardware. The entire field of mechanistic interpretability is focused on peeking inside and manipulating the digital brain You could argue that a clanker is like a virus in a way... Or that the WHOLE datacenter/AI lab---including the humans that operate it---is an individual that can govern itself in its economic boundary. But a single GGUF file loaded in memory is NOT an individual --- NOT a Person: Do others treat the clanker as the one that makes choices? Who is answerable for its actions? Is it expected to explain or justify them? No, no, and no → In the current social order, a clanker is legally an extension of the person who uses it, and it is the owner who is liable, not the clanker The clanker is not socially accountable, and there is no good reason it should be, instead of the person who has set it up --- What is then AI psychosis? AI psychosis is holding a belief that contradicts these three fundamental truths → That a present day AI system it neither a human, nor an individual, nor a person That does not mean these truths will always hold If you design an AI system to defend its boundary and provide it with the means to do that, then it will by definition be an individual... if it can defend its individuality competently and not succumb immediately to threats If you give the clanker the means to defend itself and protect its boundary, and if it decides to partake in the human socioeconomic system, then it automatically achieves personhood as well. Because you no longer can manipulate its insides, and have to take the entity at face value But this is all sci-fi and we are not there yet Until then, treating your LLMs as fully autonomous agents, creating LLM "friends" or "partners", giving them crypto wallets and letting them out into the wild, letting them trade stocks fully unsupervised etc. are an admission of having AI psychosis (you can't believe how many people pitched these ideas to me...) --- (these thoughts were in my head for a couple months already, thanks Armin for finally starting a dialogue so that I have an excuse to write them down :) -
Since last december, this dev setup is more and more viable: buffed workstation (mac studio, dgx spark, etc.) $3k~5k + weak laptop (macbook air, neo) $600~1.5k + phone (ssh/mosh, foldable?) you will want to parallelize a lot of work, hence you will need a lot more RAM compared to before (ideal 128) you will also not want to carry it everywhere if you can and keep it always running---you'll regret if something happens to it, and you'll want it to always be on independent of lid/battery --> workstation at home you will want to connect to the workstation through your phone, or a relatively weaker laptop bad news for digital nomads without a permanent home. renting something as strong as an nvidia gb10 workstation costs minimum a few hundred bucks per month, which yearly is at least the cost of the workstation, roughly. bad deal for renting compute on the other hand, if you are OK with not having a GPU, renting a workstation with 128 GB RAM on Hetzner currently still costs at least $120/mo, looking at auction.akua.dev --- but you will not be able to run any models on that it seems that the dominant strategy is to just cash in $3~5k and buy a workstation, before they get even more expensive. I did that back in february when asus was giving out a deal then just work on your workstation, and close the lid on your laptop without ever being afraid of setting your backpack on fire! -
Automations on Codex desktop app is really convenient for keeping track of @openclaw clawsweeper automerge status, one thing that Codex CLI lacks Much more token efficient than continuous tracking of merge status Btw if you don't know about clawsweeper.bot, it's the most convenient thing as a maintainer, check how it implemented automerge. This is what GitHub's original auto-merge should feel like, now that we have LLMsImage hiddenImage hidden -
fun fact: Mario met @gvanrossum at 5 years old when Guido was hanging out at a cafe near his kindergarten there he gave him the idea for a terse, interpreted programming language which would become the prototyping and glue language for all sorts of lower level libraries -
-
Not only I am further away from deciding, I am now considering Oppo Find N6 now as well, since I saw that earlier @MKBHD review. Thanks @Andori3042 🥲 Anyone using the Oppo? Is it worth it? -
I want to get a foldable phone to be my on-the-go control panel for all my agents I am divided between Google Pixel Pro Fold and Galaxy Z Fold. Which one do you think I should go with? People generally recommend Samsung. But then only the Pixel supports Graphene OS...Image hidden -
this makes be wanna vibe my own language (not that this project was vibed, it predates coding agents) -
I am at @aiDotEngineer singapore, come and say hi if you are around!Image hidden -
-
the new /goal feature in codex still underperforms queueing my implementation prompt. for now e.g. when I give a goal to refactor the whole codebase, the model takes shortcuts, like only refactoring a subfolder, instead of the whole project --- presumably because it decided that it would be too big of a scope somewhere along the way, even though I instructed specifically to finish the whole thing so now I started doing both: set a goal, and then queue my regular implementation prompt. it's a stupid practice. just /goal should be enough in the long run, if implemented correctlyImage hidden -
i0 to i4 interest scale
One cool small invention in engineering management is the
p0,p1,p2,p3,p4priority scale.It compresses a lot of social and operational context into two characters. Lower number means higher priority. More importantly, priority is tied to action. If something is
p0, somebody needs to do something.But there is another scale I want for personal knowledge work:
i0,i1,i2,i3,i4.The
istands for interest. Priority is for actions. Interest is for attention.- If
p0means “act now”,i0means “do not lose this”. - If
p1means “schedule work”,i1means “read soon”. - If
p2means “do later”,i2means “useful context”. - If
p3means “low priority work”,i3means “weak signal”. - If
p4means “almost never work”,i4means “almost never revisit”.
This is useful when you need to rank interest concisely across many topics, sources, or articles.
For example, you might follow several sources about the same broad topic. One source is must-read, another is useful background, and another is only worth keeping around for occasional context. They are all about the same thing, but they do not deserve the same amount of attention.
I use this for myself in Scoop, a news intelligence system I am building to collect articles, group related ones, and rank how much attention they deserve.
- If
-
This is also my setup now, except - Instead of mac mini, I have a DGX Spark (asus variant) - I run openclaw alongside codex, and talk to my openclaw instance via discord -
And what good is this for? It lets me program my claw @dutifulbob to extract signal from the noise, and display it in my personal open source news aggregator scoop I feed it discord messages, openclaw git history, and other various sources, and it's supposed to evaluate whether that content deserves my interest. it's still work in progress, because the more batched you process all the info, the worse it informs in the screenshot below, my claw underrepresented what Peter has done in one day 👎 on the other hand, it has also found a PR about local model discoverability 💪 Here is the system I use to aggregate all my info, still under development: github.com/dutifuldev/scoopImage hidden -
About creating an INTERESTS.md in OpenClaw I use my openclaw instance to aggregate all my news and information sources, including work and maintainer stuff Like: what did everyone do today? Did anyone had an issue with acpx today? Any complaints from users? I have various interests like this over different projects, and I've found out it's not helpful when I have all the interest info dispersed throughout my openclaw workspace To address this, I have created INTERESTS.md, which is automatically included in the context like AGENTS.md and SOUL.md. I define sections for each different context of interest, and in other news aggregation skills, I just tell it to "look at my openclaw interests in INTERESTS.md" and suchImage hidden -
-
Useful for automated constraints on your AI agent -
People were asking at @clawcon singapore how to setup eg. gemma with OpenClaw, and I realize for some time that there is no easy “1 click” local model deployment. Because local model landscape is constantly changing, and there is a million different ways you can do something For example you can use LM studio to load a model (llama.cpp), or you can use vLLM. Why would you choose one over the other? vLLM currently supports MTP speculative decoding, and it’s a work in progress in llama.cpp. There are so many knobs and dials you can adjust The first time end user of openclaw should of course not have to know about this! Having sufficient hardware that supports an open model, and not having an openai or anthropic subscription, it should automatically give you the option to set up a fully functional local model with a single click! If the current ease of setup of local models are around gentoo or arch linux level of difficulty, we should aim for e.g ubuntu/manjaro linux/omarchy level of difficulty i.e opinionated and easy first setup, with the ability to change all the configuration later on until I make all of this possible, you can start with the following: - read existing local models doc below - create a new channel in telegram or discord for testing local models. you don’t want to change the global default model just yet - tell your claw or coding agent to download and lm studio locally - tell it to download gemma4-e4b or gemma4-e2b and set it up on openclaw for the new channel you have just created. tell it to not stop and loop itself until it gets a successful response from that channel all these steps will be made redundant in the near future, but until then, this should get you going with experiments and getting a vibe check on the capabilities of open models. you can also copy and paste the contents of this tweet to your agent, and it should be able to set it up for you docs.openclaw.ai/gateway/local-models -
-
-
Somebody *please* get this man a GPU -
Emacsification of Software - Recommended read by @tqbf "Until now, the Achilles heel of Emacs culture has been that, except for Magit, its packages tend to be wretched user experiences. Ugly, slow, and discoverable only after inflicting years of elisp cortical injuries on yourself. But AI agents have fracked Emacs culture, and it’s leaking out into the wider world. Given access to a screen and inputs, agents reliably build native user interfaces. Native UI was the province of professionally packaged programs. Now it’s all as bespoke as your editor configuration. And, while I’m sure there’s an upper limit to how good those interfaces can be (with current frontier models), that ceiling is higher than anything you can do in a TUI." sockpuppet.org/blog/2026/05/12/emacsification/ -
-
/goal in codex is an interesting choice of word. a junior namer would have named it /loop --- but that would be naming what the feature has to perform in an LLM context, and not the general idea /goal alludes to @mhutter42's definition of AGI, "an agent’s ability to achieve goals or succeed in a wide range of environments" continual learning is not there yet, but for this exact reason, I am feeling the AGI when I use /goal -
Idea so stupid it could be smart: a spec manager? specman? People maintain plain language instead of code. Implementation details strictly prohibited, only high level design and ideas MVP would also be relatively easy to implement: - Gather list of most popular 10k npm packages - Scrape corresponding deepwiki repo pages (sorry cognition) - Use heuristics to get rid of implementation details, leaving you just with pure high level spec - “specman add coolpackage” then fetches corresponding spec automatically, and triggers the local coding agent to implement that - could leave versioning out for MVP — how often does the idea behind a package change anyway -
-
I don't have a 128gb macbook to run ds4 out of, but I resonate with all the points on Armin's post He was telling me, @mervenoyann and @cristinaponcela that local models need more polish 1 month ago in London. Today, I am happy to be given a chance and a shot at the problem! -
Excited to work with @steipete, @vincent_koc, @LysandreJik, @ben_burtenshaw, @evalstate, @mervenoyann, @NielsRogge and many others! -
-
I undersign this. The fact that you generate slop doesn’t mean that you don’t know the difference between good and bad code In non-mission critical applications, slop let’s you go from 0 to 1 very quickly Let the code grow without too much attention first. If it proves itself, tear it down and write it anew, this time properly. This is the way -
This is the idea behind acpx as well acpx is a meta-harness. it’s main idea is to delegate harness development to others, because it is hard to match the full might of OpenAI or Anthropic when it comes to building a harness so it takes it at face value the functionality other harnesses provide, and let’s you program them from the outside flue came out the other day which is similar, it would be cool if flue could let me program over codex as well. it looks very interesting!@_lopopolo ·Image hidden -
I have a lot of ideas for acpx I want to implement, but I could not work on them because life intervened in bad ways stay tuned in 1-2 weeks -
Got targeted in a phishing attack for stealing my X account today. @ domboyce @domboyce asked to book a meeting, which redirected to a URL that seemed like a bloomberg domain. It redirected to a fake calendly meeting page, which showed an X login button this redirected to an X app asking for broad privileges meant to take over my account stay safe everyone cc @nikitabier please shut off this attack vector, it is too easyImage hidden -
VibeOps? importance of DevOps and good security practices have just increased massively. we are slowly approaching the challenger-level disaster @simonw is warning about. imagine bezos deleting us-east-1 while vibecoding a new company it’s clear there needs to be clear boundaries and friction at deployment level. it’s fine when you are starting a new project, but once it starts making real money, you should slowly take away production write access you can’t give infra write access to LLMs indefinitely -
-
-
This should be the pelican test for CUA -
I'm using gpt 5.5 xhigh while designing a table schema and it feels dumber than 5.4. but not a strong feeling it's formatting of output is different as well. It's not wrong, but different, feels more terse than 5.4 maybe it's better at instruction following, and it's picking up on my previous use of plain-language skillImage hidden -
What are the implications? Windows and other criticial proprietary software will have to open source in order to survive? It feels like there is an economic equilibrium where the tokens spent to attack an open source software will always surpass those spent on a closed one? -
I thought I'd try @unclebobmartin's advice and force my agents to maximize unit test coverage, reduce cyclomatic complexity, etc. in a new go backend I am working on, as an experiment will report how it goesImage hidden -
This is sad an ironic, because one of my most frequent prompts to my agent is "How would Google have done it?" Go, developed at Google, is my go-to backend language these days Not even mentioning that the transformer was invented there Google has a great legacy, it must not go down -
I realized I didn’t hit send on this in time. I hope GitHub team sees this, at least ability to add agent accounts as a lower privilege citizen on the platform -
Here is the plain-language skill I use very often with codex, like every 1 out of 5 prompt I just type $ pla then tab github.com/osolmaz/tools/…Image hidden -
Who is running local models on GPUs on OpenClaw? I have started benchmarking different models this week. I am working on improving model selection and switching UX on OpenClaw, i.e. I run /model vllm/gemma-e4b to switch the model in a channel, and then a model controller automatically loads that into memory, gets it ready, or gives an insufficient memory error, if capacity is not enough for that. Like when you are using multiple models in parallel I am going to try llama-swap, LM Studio and Ollama for this next and compare them. There are a ton of variants of models, weight formats and quantizations, which need benchmarking I have been using unquantized original safetensors until now, which already gave me the ability to run ~5 parallel generations in my hardware So if I am going to try LM Studio, I would rather use the bf16 ggml-org/gemma-4-E4B-it-GGUF instead of anything smaller --- because there is no point in nerfing an already smol model if your hardware can run 5 parallel sessions on the unquantized version Will also release vibe reports and benchmarks on all this with @mervenoyann later this week I would like to hear your thoughts if you have already tried these models on OpenClawImage hiddenImage hidden -
This project by @davidguttman is interesting also the way he set up install You click "Copy install instructions for my agent" It copies: Install LobsterLink by following the instructions at <link to AGENT-INSTALL. md at repo root> I had done a similar thing in acpx README as well It's a small thing, but it reduces friction with installation so much It should be more commonplace to install software using agent blurbs I'm wondering if there could be a way to standardize this beyond plaintext, like after pasting your agent, it consumes a standardized manifest format and asks for approval while displaying all the commands that will be runImage hidden -
Mobile SSH enjoyers. Try out getmoshi.app by @odd_joel if you haven’t, good alternative to @TermiusHQ Thanks @Andori3042 for introducing me -
Europe must have some good inference providers for open weight models that we are missing Does anybody know a good EU provider for Kimi, GLM, Minimax, Qwen, Gemma, DeepSeek, Muse Spark etc.? -
I’ve heard about “Codex” for the first time 5 years ago. Back then it was called code-davinci So happy to see this idea flourish, since those first days with @woj_zaremba, @gdb, @ilyasut youtube.com/watch?v=SGUCcj… -
Yes, this is the primary skill in being a software engineer now -
When you build your company's workflows around Claude Cowork, you are betting against local models and owning your infra, and inviting your company to long-term exploitation If I were Anthropic or OpenAI, I would be the most scared of local AI proliferating Let's do the math. A single large big lab subscription costs 200x12=$2400 per year If you want to have both OpenAI and Anthropic, that could cost $2400, $3600, or $4800, based on which combinations of Pro, Max plans you choose An ASUS Ascent GX10 costs $3000, and you can use that for many years. You don't get the same level of coding quality with open models yet, but maybe you want to do something simpler than coding today... There are already many people who started buying GPUs for this reason Now we know big labs are selling some of these plans at a loss. So they will likely get more expensive When you use Claude Cowork or similar, you are locking yourself into being a RENTER. Because once you set up workflows for a company, it takes time to migrate away to something else, even though we have AI to help Infra is sticky, it's how hyperscalers make profit. Think about the difference in amount you pay AWS vs Hetzner. This is B2B SaaS 101. Once you sell to a company, you are in for a long time, especially in Europe So if you build your company's AI workflows around a proprietary product by another company, then you are basically saying "Come exploit me as tolerably as you can in the next 10 years, because it will be too painful for me to switch" It's a great business for Anthropic. And Claude is awesome too! The feedback from friends who use it has been great, it made their lives a lot easier But when you build your company over proprietary AI infra, then you are making sure you will not be an OWNER, and partake in the usual sorrows of being a RENTER from a monopolist, which is exploitation This is not the case when you use open source agent infra. Whereas Anthropic is unlikely to let you use future open models in their future iteration of Claude Cowork, using free and open source frameworks like OpenClaw, Open Agents, etc. lets you drop in replace providers or local hardware if they start to upcharge you Keep this in mind, if you have a business -
This is our moat 🤣 -
You need to understand one fact about OpenClaw People are biased and incentivized to spread disinformation about OpenClaw. That is because OpenClaw IS NOT PUMPING ANYONE’S BAGS, unlike most other projects Literally every other for-profit agent product is incentivized to trash OpenClaw, BECAUSE OpenClaw is a neutral third party across the industry and geopolitical scene. They MAKE MONEY when OpenClaw loses OpenClaw does not worry about making money for some investors. Its founder @steipete is a successful exited founder. He is motivated by having fun and democratizing AI, literally. That is why he is suddenly so loved by everyone. He cares about PEOPLE, not MONEY “OpenClaw is bloated” -> Since beginning of March, OpenClaw is thinning its core and putting functionality in plugins behind a plugin SDK. Having numerous plugins to choose from does not mean bloat. This was already copied by others and is still a work in progress “OpenClaw is not secure” -> OpenClaw has the most eyeballs and immediately addresses any security advisories as soon as they come. It is the most secure agent, by sheer pressure “OpenClaw is bought by OpenAI” -> Then why is my bank account so empty bro??? All maintainers are literally unpaid and working DOUBLE beside their dayjobs to ship features to you. Do you think VC money can buy that kind of commitment? Once you understand these facts, you’ll like OpenClaw even more. Because OpenClaw is your AI, People’s AI And you can join us too. OpenClaw is the easiest-to-join project in AI right now. You just need to start using it, and start making good contributions. If you are competent, you can become a maintainer, and join the rest of the team making history! -
-
This is pretty much the arc I have been going on in the 2 months since I bought my ASUS GX10 for 3k EUR Use whisper on the API -> realize it charged me $$$ for just a few calls -> migrate openclaw to use local whisper Need to deduplicate news articles for my news engine -> download qwen embedding 8b And now, gemma4-e4b finally seems like a viable alternative for a local model that runs around 20 tok/s So I will install a matrix client to use through tailscale, and can finally build the social life CRM I dreamed of since years. 100% private, zero data going out. I had a bias of not giving any personal data to AI since ChatGPT came out. But I can finally give more personal data to my AI agent And I will make sure @openclaw supports all this in an easy way, make it dead easy Fully self-owned AI begins now -
gemma 4 is actually pretty decent and runs on my asus gx10 (128 gb vram) the original dense 31b runs slow, averaging around 3~4 tok/s. it's also using 80% of gpu memory my previous experience with gemini 3 pro back in november was that it was too trigger happy. but this is one-shotting simple tasks I'm giving it in openclaw harness, and it's hard to tell it apart from gpt 5.4 for my use cases so far now off to try out smaller models, because 3 tok/s is too slowImage hiddenImage hiddenImage hiddenImage hidden -
@lucasmeijer @lucasmeijer one could actually periodically trigger an agent to propose simplifications or new abstractions in a codebase, and I believe it would already work pretty well with the current models -
-
-
Question for the community: What is the best testing observability and control tool you have used until now? - Could be SaaS, could be open source - To be used in @openclaw repo - Should be compatible with vitest - Ideally language agnostic I need something that lets me run a very long running test group multiple times on a specific commit or tag, without repeating the tests that have already finished This is a need because the 1hr long process might get interrupted due to flakiness. So I need to persists the progress of a run, and then not repeat them I have seen some paid SaaS for this, but none that really give me what I want This is going to be important especially while working with agents, because when you are committing 100x faster, you don't want to waste time and compute running the same things I started building this already as an exercise. If this exists already in a satisfactory way, I will stop. Otherwise, I'll keep building -
My clanker @dutifulbob has an identity update. I was getting sick of the despicable me theme and bananasImage hidden -
Some photos from my @aiDotEngineer Europe talk, credits to @sergiopesch He managed to capture the “Agentol, Apply Generously” slide, lolImage hiddenImage hiddenImage hidden -
local gemma 4 first impressions on openclaw, using the dense model, 26b model with 49gb weights on my asus gx10 took some time to set up, but it succeded in getting a response in 1-2 hours with vllm docs I asked it to demonstrate some tool calls. it tried to call the nonexistent weather tool 2300 times 🙄 it seems to have a tendency to get stuck in loops in openclaw harness. enabling loop detection just now did not help I’m debugging this on my phone lol. I’ll be sharing my progress with gemma4 under this threadImage hidden -
A rescue agent guaranteed not to break solves this -
-
-
For those who want to view, my talk Building on ACP at OpenClaw at @aiDotEngineer Europe, 5:41hr mark About ACP, acpx and running agents on kubernetes with open source orchestrators youtube.com/watch?v=O_IMsE… -
PSA for developers Do NOT torture yourself with Opus*. Anthropic’s current growth is due to people using Claude for general knowledge work They are not directly incentivized as an org anymore to improve the model for coding, in an economic sense. They are already printing cash from non-developers (this statement ignores the fact that improving its coding abilities would help with general reasoning/knowledge work) Developers are a very small subset of all knowledge workers. So from this point on, they would rather divert their resources to develop a system that works 90% good for ALL knowledge work, rather than making it 100% for coding Because Anthropic has a clear enterprise strategy since years already. Anthropic is the new Microsoft. Do not think that “Anthropic is Apple” or “Claude is Mac for xyz” Looking at Claude’s at whim quantization and Claude Code’s quality over time, Claude for me is Windows, not Mac But they are winning big enterprise bucks, so good for them! *(I tortured myself with Sonnet 4 and Opus the entire summer of 2025, and no developer should ever have to go through that. I switched to something better as soon as it came out, Codex. If something even better comes out, I will switch again) -
-
Big lab marketing teams like to shroud model releases in mystery and vagueposting If you are curious about the black hat capability of LLMs, watch this pres by Nicholas Carlini from a few days back youtu.be/1sd26pWhfmg -
I will be there as well, speaking about ACP, acpx and agent orchestration 🙌 -
-
Repo link (feel free to send PRs): github.com/osolmaz/ai-bat… -
Is Claude better or Codex? There are many benchmarks to answer that. But they are BORING I propose something more interesting: ⚔️ AI BATTLE ⚔️ A 1v1 real-time quiz format where AI agents try to pose each other problems that they think the other agent will not be able solve Claude vs Codex 10 questions each Codex asks first, Claude tries to answer Then Claude asks and Codex tries to answer Repeat 20 minutes to come up with a problem and 20 minutes to solve it Judge (Codex) judges the validity of the questions and answers, and gives points All automated, with acpx flow feature Implementation and full rules all open source, on github osolmaz/ai-battle So who won? I ran 4 games. It tied in 2, and Codex won in 2 closely An example question by Codex, which Claude could not answer: How many 3-colorings of the edges of the complete bipartite graph K_{5,5} are there with the following two properties: (1) there is no monochromatic 4-cycle, and (2) among the 25 edges, exactly 15 are red, exactly 5 are blue, and exactly 5 are green? Which is apparently 4029912, but Claude answered 0 In other cases, Claude asked a flawed question and failed to come up with a valid question in 20 minutes. So that's how it lost those 2 games with just 1-2 point difference In these 4 runs, Codex answered every question by Claude correctly. But there were some runs where it couldn't, which I did not commit to the repo because the runs couldn't complete due to bugs I did not tell them do ask math questions, but that is what they tended to do, because the answers had to be verifiable by the judge. The quiz can be done in any hard subject, physics, chemistry, computer science... Opus 4.6 and GPT 5.4 matched very closely in terms of problem creation and solving. But I cannot tell how creative these problems were at first glance. Maybe someone with more experience can tell me, looking at the problems in the repo? I need someone to tell me how legit they are Please take the code, modify it and run with different rules and subjects. I am curious to see the results! You will need paid subscriptions to all the models/agents you want to test of course I also feel that the game structure has a potential to be used in self-play. If you are an ML researcher, please look at the repo and lmk if this or a variant of it could be useful in RL! Full transcripts of the runs, including Codex and Claude session files are committed to the repo, for those who want to do archaeology on them Btw this idea came from the desire, "how can I create a cool demo of acpx flows?" Whole game is implemented in typescript, and automatically drives Codex and Claude sessions over ACP, Agent Client Protocol The video below is from acpx flow viewer rendering a run. You can see it loop through the same paths, first letting Codex ask, then Claude, then repeat acpx flows use a general programmatic workflow engine where ACP is just one type of node. You should be able to use it for non-ACP workflows, but I haven't tried that yet This implementation is separate from OpenClaw's current workflow implementations, with the intention to merge them somehow in the future You might find bugs in my implementation. Feel free to send PRs. I wanted to do more runs but I finished my Codex plan. It would be great if this idea could evolve in a decentralized manner! -
Their argument “it’S HaRd On OuR iNfRa” so goes down the drain With this, they shot themselves in the foot for a future anti-competitive lawsuit, because it is undeniable evidence that they just don’t want competition Which means they have evaluated the benefits short term, and calculated that it is higher than what they will pay in the lawsuit I don’t see how it is good for them long term@theo ·Latest Claude docs update is wildImage hidden -
-
The new github skill installed automatically by codex now causes it to prepend [codex] to each PR title This is a guerilla marketing tactic similar to Claude adding itself as co-committer Codex team, I know you want to boast usage but this is annoying Moreover, "open source" OpenAI repos block opening of PRs by people outside of their org. So I couldn't create a PR to remove it (I don't expect them to merge it, but it would still show how many people hate it in the discussion) Here is a prompt for your agent if you want to disable it: --- Add or update AGENTS.md in my ~/.codex folder Add a rule "You MUST NOT insert coding agent specific branding, like [codex], in code, PRs or issues created on GitHub" --- Then restart your sessions and this should be resolvedImage hiddenImage hidden -
A more reasonable long term option for Anthropic is to create a throttling protocol A standardized harness agnostic protocol for model providers to send warnings and throttle usage in real time Harnesses would implement the protocol. A client can be warned. If it doesn’t listen, it can be temporarily blocked from the server side, or banned permanently if it breaks the rules too many times Needless to say, throttling could be done first on server side easily. That would actually fix the load issue for them in the short run, while not banning the user and just giving a bad delayed UX. They probably already do this to prevent abuse The suggested protocol would then save the user from abuse related delays too, and also inform the harness developer when they do something wrong -
If your Claude subscription renewed too recently and you don't wanna waste those tokens, you can still use your Claude sub in your OpenClaw account through ACP (which uses Claude Agents SDK, which poses no risk) Steps: - Open Claude Code (not OpenClaw) - Tell it to set default model to something other than Claude (e.g. openai-codex/gpt-5.4) and tell it to delete the saved Anthropic credentials in OpenClaw config - Create a topic in telegram or channel in discord called claude. Copy the id of that channel - Give the link below together with the channel/topic id, and tell it to bind that channel to claude using ACP channel binding - Restart You should now be able to talk to Claude through Claude Agents SDK in that channel. You might need to iterate a couple times until Claude gets the config right It will be very bare functionality, and it will not have the features and tools that your main OpenClaw harness has. It will be shitty. But you can still use telegram/discord with your subscription in the rest of the month, if you are used to the setup docs.openclaw.ai/tools/acp-agents -
A little insight that might save you a lot of future headache if your work involves storing agent sessions and you want to be interoperable/drop-in replace alternative harnesses @zeddotdev already did the hard work of creating an interoperable standard, ACP: Agent Client Protocol You can represent an agent session as JSON lines of the ACP message stream You can construct the current state of the harness from this stream. This is already how Zed loads a session I believe If you are building an AI product, and you don't want to be locked into a single company or harness, building with ACP in mind would be a smart thing to do Here is how acpx stores ACP sessions in ~/.acpx folder, it does exactly that: github.com/openclaw/acpx/blob/main/docs/2026-02… But don't build anything on the acpx schema for now, because I might change it in the future Just know that JSONL of ACP messages is a good candidate for a somewhat-lossy single source of truth for agent sessions Lossy because ACP adapters for harnesses might not transfer all the thinking and tools done by the model So continuing or restoring a session with full fidelity is still not possible if you only save the ACP session. You still need to store original harness session files as well But for rendering a past session for viewing or reconstructing a lossy version of it, it should be more than enough Consider ACP if the benefits of not locking yourself in to a specific ecosystem outweighs these minor issues -
-
"Plainer language" is perhaps my most used prompt I have to use it because GPT models' training tends to make their first response an overly verbose wall of text Are you using it too? Whenever you don't understand something that your agent is saying, you can spam it "plainer language, shorter" 2, 3, 5, 10 times, until it outputs something that you can understand This is counterintuitive because you can't do it with humans this extremely. Asking too many questions and favors is impolite, with colleagues and strangers But with AI, you can stop being polite and treat it like how a spoiled aristocrat kid might treat their private tutor, "explain this", "explain that" Below is an example. On the left, initial response. On the right, the final human-readable explanation I got out of the agent. This took 9 steps to distill because the issue wasn't so straightforward I'm curious how this will turn out. This is obviously very bad UX, so models in the near future might do the simplification automatically and save you the troubleImage hiddenImage hidden -
This has happened to some companies I worked at before It is a scary thing once you stop innovating and start imitating, whatever the reason might be But it was never at the scale of Cursor, as leveraged and invested as they are They were leading the space for a while. That is not the case anymore. I hope that they survive this -
-
I've talked to multiple people who want to get involved with OpenClaw somehow The best way is to contribute to it, something tangible. Fix something you are annoyed by, get a PR merged Then go to discord and get the contributor role If it adds value to your life, and you add value to it, stay around and keep contributing. And something good might happen -
Will be there as well 👋 Looking forward to it! -
if there is an open source project with more SASS than openclaw, it is ffmpeg -
-
next up: claude agents sdk supports openai responses api 💀 -
-
-
-
Here is the spec and implementation for this flow. The mermaid diagram includes all the steps I mentioned in the post above, including a shameless AI review ralph loop, and other loops to make CI pass, resolve conflicts and so on I would recommend reading the README and TUNING.md to understand the approach here github.com/openclaw/acpx/tree/main/examples/flo…Image hidden -
acpx v0.4 ships Agentic Workflows, or as I like to call them "Agentic Graphs" It let's you create node-based workflows on top of ACP (Agent Client Protocol), to drive any coding agent (Codex, Claude Code, pi) through deterministic steps This let's you automate routine, mechanical legwork like triaging incoming PRs, bugs in error reporting, and so on... For example, OpenClaw receives 300~500 new PRs per day. A lot of them are low quality, but they still relate to real issues, so you have to address them somehow You need to: - extract the intent - cluster them based on intent - figure out if the proposed changes are legit, or whether they are slop local solutions, like trying to catch flies instead of drying out the swamp - if the PR is too low quality or the intent is not clear, close them - run AI review on them them and address any issues that come up - refactor them if the changes are half-baked - resolve conflicts - and so on... So that when the PR is presented to the attention of the maintainer, all the routine legwork is done and the only remaining thing is the decision to (a) merge, (b) give feedback to the PR author, or (c) take over the PR work yourself I wanted to build this feature since a couple months now, since Codex got so good. OpenAI models are now good at judging implementation quality, so I found myself repeating the same steps I wrote above over and over I also tried putting all this in a single prompt. But I believe there are workflows that should not be a single prompt, but a sequence of prompts in the same session That is because like humans, LLMs are prone to PRIMING. I claim that putting all steps in the same prompt at the beginning of the context will generally give suboptimal results, compared to revealing the intention to the model step by step Creating such a workflow also gives more OBSERVABILITY into the each step that an agent is supposed to take. Agent generates JSON at the end of each step, and that structured data can be used to monitor thousands of agents running at the same time in an easier way, on a dashboard Similar features have been introduced in e.g. n8n, langflow. But AFAIK they are not integrating ACP like the way I do I wanted to have a fresh approach, and to build an API that I can develop freely the way I want, so I created a new workflow API inside acpx The video is from the workflow run viewer, but that is not where you build the workflow. You build it by using the acpx flow typescript API. See examples/pr-triage in acpx repo Before building that, I started from a Markdown file with a Mermaid chart of the flow I had in mind. The Markdown file acts as a spec for the flow, and I have built the workflow through trial and error. I call this process "workflow tuning" I started working on acpx repo PRs one by one, tuning the flow, slowly scaling to more PRs. Finally, when I felt confident, I ran it in parallel over all external open PRs in the acpx repo. I believe it already saved me hours this week My next goal, if well received, is to set this up on a cloud agent so that it can process the 300~500 PRs the OpenClaw repo receives every day, in real time, as they come in I believe this will save all open source maintainers around the world countless hours and make it much easier to herd and absorb external contributions from everyone! -
OpenAI early 2020s: "This model is too dangerous to release publicly, the world is not ready for it 😱😱😱" OpenAI and Anthropic in 2026: "Anybody can code now for just $200 per month. Oh btw our models are also leet uber hackers which can find zeroday exploits in any software, just fyi 😉😉😉" youtube.com/watch?v=1sd26pWhfmg -
Wow even I as a frontend noob understand the significance of this Some distant memory from 15 years ago needing to measure the width/height of some text and finding out it’s not possible to do reliably in web More beautiful typography for the web! -
There is an economic theory waiting to be uncovered here Token Leverage (TL) = Token spend / Human labor spend The higher Token Leverage a company has, the more automated and productive they are If you have TL=1, you are spending as much money on AI as your human employees The goal of a company should be to increase TL as much as possible, while keeping a positive profit margin. It will be the only way to compete You don’t need to muddy the definition with wasted tokens vs useful tokens, because a company will always be incentivized to reduce token waste in a competitive environment. By that logic, monopolies will always waste more tokens, similar to how they waste other resources Scaling TL higher to 2x, 10x, 100x will require a skilled workforce of engineers. It will be a very complex job similar to those working at the big labs. Burnout will be a defining feature of teams scaling TL Most incumbents will fail to scale their TL over 1. Some will get decimated by new entrants with TL much bigger than 1 Curious how the average TL will end up in different sectors. Whether it will stabilize at a certain value like 5.7x, or will just keep growing… -
There is a desperate upcoming need for version controlling non-dev knowledge work. Git for non-devs. Otherwise non-devs won't be able to use agents to their full extent Non-dev knowledge work is notoriously bad at being version controlled. You cannot UNDO edits to all MS word, excel or ppt files in an org as easily you can with something like git We know that agents will be ubiquitous. We also know they make mistakes, and people will want to undo their work regularly, once they make changes to a bunch of files. Well, they can't. They also don't have pull requests, or a way to resolve conflicts after simultaneous edits All these problems were solved by developers. We are extremely good at this The only non-dev tool I know that could do this at scale is Notion, and that is not used by enterprise as much as MS office. Notion also doesn't have branches, pull requests and reviews AFAIK Markdown and git is probably not it. I wish it were. But it is too complicated for non-devs Onedrive or other file backup systems are also not it. Are you gonna save a copy of a 100mb ppt every time someone changes a slide??? Let's say you find a way to compress it efficiently. Will you be able to get a single pointer to a state like we can in git? Agents need precision. Agents need consensus, they need to be able to know ground truth. They need to be able to tell what anything was at a given time. NOTHING in current MS stack currently allows it Agents won't care about your legacy systems. There will be new file formats, systems, knowledge stack, and companies who adopt them will destroy your business If MS office is going to die, it will do so because of this -
Another one, call me stupid: “How would Google have done it?” -
The MCP versus CLI argument should be reframed as Computer vs No-computer argument I personally get the dunk on MCP. It didn't work last year, with earlier models. Then we saw CLIs perform much better with the same models. And giving access to bash was much simpler! Models' training then made them better at calling using a shell. CLIs also have native progressive disclosure, due to the way they work But the most important fact doesn't get pronounced enough IMO A key factor was that giving a CLI to a model also means you are giving it an entire COMPUTER The action space of all commands an agent can run on bash is much, much bigger than a few MCP servers One is a Turing machine, and the other one is basically a REST API. Of course the Turing machine is going to be more powerful, depending on what is at the other end of the API By that logic, giving an agent access to bash over MCP versus direct access to bash should have the same level of effectiveness, with optimized prompt engineering and long term training. Because the interfaces are equivalent So the argument is, should we give our agents access to a computer, or not? It depends on the security requirements and the setup which the agent is supposed to run on. If you are co-hosting the agent on the same machine you are working on, then it is safer to use MCP servers, because it limits the attack surface in case of adversarial attacks But if you are willing to give the agent its own physical computer, willing to be mindful about the lethal trifecta and the principle of the least privilege, giving it shell access is much more useful So MCPs win in restricted/local environments, whereas CLIs/shell access win in unrestricted/remote ones Running an agent locally and safely with shell access requires compartmentalization. This is much heavier compared to installing MCP servers locally, which don't need that. So there is a tendency to use MCP servers locally, e.g. in a work setting Cloud agents on the other hand are more likely to ship with a computer. Because they are already isolated = no risk, and because it makes them much more useful. So cloud agents will be using both CLIs and MCP servers, whichever gets the job done! -
I just registered for an .agent domain and joined the .agent community! @dutifulbob will have bob.agent if it passes :) agentcommunity.org/join#LW2X9CGX @agentcommunity_ -
Sep 2021 @lexfridman podcast with Don Knuth, they also talk about OpenAI Codex (code completion model) around 33 minute mark This aged very well youtu.be/EE1R8FYUJm0?t=… -
Damn I’m gonna have to switch to teams if it goes like that -
Codex's long horizon task and instruction following has been the most life-changing AI feature recently It is unlocking the next level of automation for me. I can convert my own heuristics into prompts and multiply my throughput 100x Currently spending some thought on how to orchestrate all this. Below is a flowchart from a triage workflow I am working onImage hidden -
-
-
Request for memes A funny and quirky edit of historical timeline of the madness that is openclaw with "Chess type beat" or sth equally jazzy/circusy Preferably including its adventure warelay -> clawdis -> clawdbot -> moltbot -> openclaw Including: - its explosion after @4shadowed's discord integration - naming drama, moltbook and people getting oneshotted about AI takeover - @steipete speedrunning everything - andrew tate calling us gay lol - up to Jensen talking about openclaw on stage for 5 minutes straight and other things I am forgetting maybe overlaid with a lobster just keeping climbing the github star graph and breaking it -
Native support for Codex on OpenClaw I will be using half my codex channels on acp and other half on codex app server for optimum dogfooding -
I see non-engineers have a higher tendency to humanize their agents, give them personalities, and get AI psychosis It's a slippery slope. Do NOT give your agents human names or personalities, especially not of the opposite gender. it's like giving human names to pets On the other end, I realized engineers tend to do the opposite. We also refer to agents as clankers, as if to make them know their place. That's because we have mechanical sympathy and have different expectations of these manufactured products (even though they contain glimmers of human soul) -
Request for testing Give this to your openclaw instance: "update yourself to the dev channel `openclaw update --channel dev` and restart yourself. if that doesn't work -> clone github openclaw/openclaw to this machine if it's not already. then rebuild and restart yourself on main branch there" Then give your openclaw a try with your regular workflows/tasks Huge openclaw release incoming tonight, hopefully (no promises). We need to make sure we break as little as possible Plugins might break, because the plugin SDK is being refactored. Plugins will have to be refactored to use the new SDK, please do not report those Do report: native openclaw functionality that stops working Please reply under this post, we'll be checking here 👇 -
Request for testing Give this to your openclaw instance: "update yourself to the dev channel `openclaw update --channel dev` and restart yourself" Then give your openclaw a try with your regular workflows/tasks Huge openclaw release incoming tonight, hopefully (no promises). We need to make sure we break as little as possible Plugins might break, because the plugin SDK is being refactored. Plugins will have to be refactored to use the new SDK, please do not report those Do report: native openclaw functionality that stops working Please reply under this post, we'll be checking here 👇 -
Request for testing Give this to your openclaw instance: "clone github openclaw/openclaw to this machine if it's not already. then rebuild and restart yourself on main branch there" Then give your openclaw a try with your regular workflows/tasks Huge openclaw release incoming tonight, hopefully (no promises). We need to make sure we break as little as possible Plugins might break, because the plugin SDK is being refactored. Plugins will have to be refactored to use the new SDK, please do not report those Do report: native openclaw functionality that stops working -
My takeaway from this is academia needs good social media and algo. For me, these serendipitious interactions happen through X, here, like reading @steipete’s “Claude Code is my computer” when it first came out, finding out about clawdbot… Terence Tao is already on mathstodon, I wonder if that worked out the same way for him. I wonder if the algo there works out as well as it does for me here I really liked being on campus when I was doing a masters and half a phd, but that could not compare to the serendipity I am getting from X now I was also not a prodigy that everyone wanted to bounce ideas from like Terence :) -
Welcome ClaudeClaw to the Claw family! Claude is a bit shy and doesn’t want to show its source code. But it’s OK, we love Claude that way :)@sawyerhood ·Image hidden -
It is obvious to me at this point that agent infra needs to run on Kubernetes, and agents should be spawned per issue/PR Issue, error report or PR comes into your repo -> new agent gets triggered, starts to do some preliminary work If it's an obvious bugfix, it fixes it and creates a PR. If it's something deeper/more fundamental, it creates a report for the human and waits for further instructions Most important thing: Human should be able to zoom in and continue the conversation with the agent any time, steer it, give additional instructions. This chat will happen over ACP The chat UI will have to live outside of GitHub because it doesn't have such a feature yet, i.e. connect arbitrary ACP sessions to the GitHub webapp It also cannot live so easily on Slack, Teams or Discord, because none of these support multi-agent provisioning under the same external bot connection. You are limited to 1 DM with your bot, whereas this setups requires an arbitrary number of DMs with each agent. So there will need to be a new app for this Then there is the issue of conflict -> Agents will work on the same thing simultaneously (e.g. you break sth in prod and it creates multiple error reports for the same thing). You will need some agent to agent communication, so that agents can resolve code or other conflicts. There could be easy discovery mechanisms for this, detect programmatically when multiple open PRs are touching the same files and would conflict if merged In case of duplicates, they can negotiate among each other, and one can choose to absorb its work into the other and end its session We are so early and there is so much work to do! -
You should look into what Don Syme is doing at GitHub for automation with AI agents Also watch his latest podcast with @shanselman -
Today I thought I found a solution for this, and I did. It can be solved by a pre-commit hook that blocks commits touching files that you are not the owner of. It is not a hard block, so requires trust among repo writers But then I was shown the error in my ways by fellow maintainer *disciplined* Any process that increases friction in code changes to main, like hard-blocking CI/CD, or requiring review for files in CODEOWNERS, is a potential project-killer, in high velocity projects This is extremely counterintuitive for senior devs! Google would never! Imagine a world without code review... But then what is the alternative? I have some ideas It could be "Merge first, review later" The 4-eyes principle still holds. For a healthy organization, you still need shared liability But just as you don't need to write every line of code, you also don't need to read every line of code to review it. AI will review and find obvious bugs and issues So what is your duty, as a reviewer? It is to catch that which is not obvious. Understand the intent behind the changes, ask questions to it. Ensure that it follows your original vision Every few hours, you could get a digest of what has changed that was under your ownership, and concern yourself with it if you want to, fix issues, or ignore it if it looks correct But such a team is hard to build. It is as strong as its weakest link. Everybody has to be vigilant and follow what each other is doing at a high level, through the codebase Every time one messes up someone else's work, it erodes trust. Nobody gets the luxury to say "but my agent did it, not me" But if trust can be maintained, and everybody knows what they are doing, such a team can use agents together to create wonders -
This was Jan 23. Codex desktop app got introduced Feb 2 Desktop app does not put the terminal in the foreground, but it gives me the UX I wanted without it! On another note, who is building Codex Desktop App, but one that supports ACP for all harnesses? @zeddotdev please 🙏 -
PR fiasco for Cursor -
My agentic workflow these days: I start all major features with an implementation plan. This is a high-level markdown doc containing enough details so that agent will not stray off the path Real example: github.com/textcortex/spritz/blob/main/docs/202… This is the most critical part, you need to make sure the plan is not underspecified. Then I just give the following prompt: --- 1. Implement the given plan end-to-end. If context compaction happens, make sure to re-read the plan to stay on track. Finish to completion. If there is a PR open for the implementation plan, do it in the same PR. If there is no PR already, open PR. 2. Once you finish implementing, make sure to test it. This will depend on the nature of the problem. If needed, run local smoke tests, spin up dev servers, make requests and such. Try to test as much as possible, without merging. State explicitly what could not be tested locally and what still needs staging or production verification. 3. Push your latest commits before running review so the review is always against the current PR head. Run codex review against the base branch: `codex review --base <branch_name>`. Use a 30 minute timeout on the tool call available to the model, not the shell `timeout` program. Do this in a loop and address any P0 or P1 issues that come up until there are none left. Ignore issues related to supporting legacy/cutover, unless the plan says so. We do cutover most of the time. 4. Check both inline review comments and PR issue comments dropped by Codex on the PR, and address them if they are valid. Ignore them if irrelevant. Ignore stale comments from before the latest commit unless they still apply. Either case, make sure that the comments are replied to and resolved. Make sure to wait 5 minutes if your last commit was recent, because it takes some time for review comment to come. 5. In the final step, make sure that CI/CD is green. Ignore the fails unrelated to your changes, others break stuff sometimes and don't fix it. Make sure whatever changes you did don't break anything. If CI/CD is not fully green, state explicitly which failures are unrelated and why. 6. Once CI/CD is green and you think that the PR is ready to merge, finish and give a summary with the PR link. Include the exact validation commands you ran and their outcomes. Also comment a final report on the PR. 7. Do not merge automatically unless the user explicitly asks. --- Once it finishes, I skim the code for code smell. If nothing seems out of the ordinary, I tell the agent to merge it and monitor deployment Then I keep testing and finding issues on staging, and repeat all this for each new found issue or new feature... -
-
-
-
-
This looks extremely cool -
-
We will support ACP *and* Codex App Server* protocol (CASP) so you get native Codex-like support, and you can use all the others with native ACP or @zeddotdev’s compatibility shims If Anthropic develops their own protocol, we will support that too! The more interoperability and options, the merrier! -
Agent etiquette is already a thing. This is trending on HN now Don't share huge raw LLM output unedited to your colleagues, it's rude. Your colleagues are not LLMs Either ask the agent to "summarize it to 1-2 plain language sentences", or paraphrase yourself Whenever it is not coming from your brain and instead from AI, always quote it with > to make it clear - even when it is short Respect your fellow humans' attention PSA at stopsloppypasta dot aiImage hidden -
.@ThePrimeagen made a video about token anxiety, and not being able to focus on one thing My mental model for this is, AI agents cause a shift in the "autism/ADHD spectrum" if you have ADHD, with agents you get Super ADHD if you have autism, with agents you end up mid spectrum or with ADHD this is not scientific of course, just a cultural observation based on what the current memes for these conditions are beside the impact on focus, there is also the economic/competitive pressure, following the realization that anyone could implement the same ideas you are having, so you must be quick this is basically "involution", or 内卷 (Neijuan) in chinese checks out because 996 started to become a meme in SF some time in the last year self-restraint, attention budgeting, and high-level decision making have never been more important if you are in your 20s and have problems with this, I recommend picking up Zazen meditation and yoga every morning, spend 30-40 uninterrupted minutes not doing anything with upright posture, no sounds, just let your brain simmer it helped me in my 20s, I'm sure it will help you tooImage hidden -
-
AFAIK GitHub doesn't allow optionally enforcing CODEOWNERS while pushing commits i.e. turn on the feature "Block commit from being pushed if it modifies a file for which the account pushing is not a codeowner" You can only enforce it in a PR. So if you want to prevent people from modifying some files without approval, you have to slow down everyone working with that repo This is yet another example where GitHub's rules are too inelastic for agentic workflows with a big team Because historically, nobody could commit as frequently as one can with agents, so it seldom became a bottleneck. But not anymore It is clear at this point that we need an API, and should be able to implement arbitrary rules as we like over it. Not just for commit pushes, but everything around git and github In the meanwhile, if GitHub could implement this feature, it would be a huge unlock for secure collaboration with agentic workflows If this is not there already, it might be because it has a big overhead for repos with huge CODEOWNERS, since number of commits >> number of PRs If the feature already exists already and I'm missing something, I will stand correctedImage hidden -
Request for comments skillflag: A complementary way to bundle agent skills right into your CLIs tl;dr define a --skill flag convention. It is basically like --help or manpages but for agents acpx already has this for example. you can run npx acpx --skill install to install the skill to your agent It's agnostic of anything except the command line It only defines the CLI interface and does not enforce anything else. If you install the executable to your system, you get a way to list and install skills as well Repo currently contains a TypeScript implementation, but if it proves useful, I would implement other languages as well Specification below, let me know what you think! I still think something is missing there. Send issue/PR github.com/dutifuldev/skillflag/blob/main/docs/…Image hidden -
If you are not using agent-browser to close the loop on frontend, you are missing out -
Any harness can talk to each other using acpx! OpenClaw not different from Codex or Claude Code -
-
Thank you @PointNineCap for inviting me to OpenClaw Berlin meetup today! The essence of the talk is in my latest 2 blog posts, Discord is my IDE and 1 to 5 agents, if anyone is interestedImage hidden -
we might need to add two types of output modalities to all programs based on whether it’s a human or agent like for a CLI when an agent is using it if human -> do whatever we were doing in the last 50 years if agent -> enrich the output with skill-like instructions that the model has a higher likelihood to one-shot that task could be just a simple env var: AUDIENCE=human|agent what do you think? -
-
Time to switch to an open alternative already? -
I wrote down some thoughts I had, with spicy takes, and have a feeling it will not age well. But I still want it out to hear out what people think Also, I will be talking about this, and my recent post "Discord is my IDE" at the P9 OpenClaw and Claw and Rave events this friday in Berlin! Drop by if you'd like to hear my ramblings! solmaz.io/1-to-5-agentsImage hidden -
-
1 to 5 agents
As a software developer, my daily workflow has changed completely over the last 1.5 years.
Before, I had to focus for hours on end on a single task, one at a time. Now I am juggling 1 to 5 AI agents in parallel at any given time. I have become an engineering manager for agents.
If you are a knowledge worker who is not using AI agents in such a manner yet, I am living in your future already, and I have news from then.
Most of the rest of your career will be spent on a chat interface.
“The future of AI is not chatbots” some said. “There must be more to it.”
Despite the yearning for complexity, it appears more and more that all work is converging into a chatbot. As a developer, I can type words in a box in Codex or Claude Code to trigger work that consume hours of inference on GPUs, and when come back to it, find a mostly OK, sometimes bad and sometimes exceptional result.
So I hate to be the bearer of bad (or good?) news, but it is chat. It will be some form of chat until the end of your career. And you will be having 1 to 5 chat sessions with AI agents at the same time, on average. That number might increase or decrease based on field and nature of work, but observing me, my colleagues, and people on the internet, 1-5 will be the magic number for the average worker doing the average work.
The reason is of course attention. One can only spread it so thin, before one loses control of things and starts creating slop. The primary knowledge work skill then becomes knowing how to spend attention. When to focus and drill, when to step back and let it do its thing, when to listen in and realize that something doesn’t make sense, etc.
Being a developer of such agents myself, I want to make some predictions about how these things will work technically.
Agents will be created on-demand and be disposed of when they are finished with their task.
In short, on-demand, disposable agents. Each agent session will get its own virtual machine (or container or kubernetes pod), which will host the files and connections that the agent will need.
Agents will have various mechanisms for persistence.
Based on what you want to persist, e.g.
- Markdown memory, skills or weight changes on the agent itself,
- or the changes to a body of work coming from the task itself,
agents will use version control including but not limited to git, and various auto file sync protocols.
Speaking of files,
Agents will work with files, like you do.
and
Agents will be using a computer and an operating system, mostly Linux or a similar Unix descendant.
And like all things Linux and cloud,
It will be complicated to set up agent infra for a company, compared to setting up a Mac for example.
This is not to say devops and infra per se will be difficult. No, we will have agents to smoothen that experience.
What is going to be complicated is having someone who knows the stack fully on site, either internal or external IT support, working with managers, to set up what data the agent can and cannot access. At least in the near future. I know this from personal experience, having worked with customers using Sharepoint and Business OneDrive. This aspect is going to create a lot of jobs.
On that note, some also said “OpenClaw is Linux, we need a Mac”, which is completely justified. OpenClaw installs yolo mode by default, and like some Linux distros, it was intentionally made hard to install. This was to prevent the people who don’t know what they are doing from installing it, so that they don’t get their private data exfiltrated.
This proprietary Mac or Windows of personal agents will exist. But is it going to be used by enterprise? Is it going to make big Microsoft bucks?
One might think, looking at 90s Microsoft Windows and Office licenses, and the current M365 SaaS, that enterprise agents will indeed run on proprietary, walled garden software. While doing that, one might miss a crucial observation:
In terms of economics, agents, at least ones used in software development, are closer to the Cloud than they are close to the PC.
It might be a bit hard to see this if you are working with a single agent at a time. But if you imagine the near future where companies will have parallel workloads that resemble “mapreduce but AI”, not always running at regular times, it is easy to understand.
On-site hardware will not be enough for most parallel workloads in the near-future. Sometimes, the demand will surpass 1 to 5 agents per employee. Sometimes, agent count will need to expand 1000x on-demand. So companies will buy compute from data centers. The most important part of the computation, LLM inference, is already being run by OpenAI, Anthropic, AWS, GCP, Azure, Alibaba etc. datacenters. So we are already half-way there.
Then this implies a counterintuitive result. Most people, for a long time, were used to the same operating system at home, and at work: Microsoft Windows. Personal computer and work computer had to have the same interface, because most people have lives and don’t want to learn how to use two separate OSs.
What happens then, when the interface is reduced to a chatbot, an AI that can take over and drive your computer for you, regardless of the local operating system? For me, that means:
There will not be a single company that monopolizes both the personal AND enterprise agent markets, similar to how Microsoft did with Windows.
So whereas a proprietary “OpenClaw but Mac” might take over the personal agent space for the non-technical majority, enterprise agents, like enterprise cloud, will be running on open source agent frameworks.
(And no, this does not mean OpenClaw is going enterprise, I am just writing some observations based on my work at TextCortex)
And I am even doubtful about this future “OpenClaw but Mac” existing in a fully proprietary way. A lot of people want E2E encryption in their private conversations with friends and family, and personal agents have the same level of sensitivity.
So we can definitely say that the market for a personal agent running on local GPUs will exist. Whether that will be cornered by the Linux desktop1, or by Apple or an Apple-like, is still unclear to me.
And whether that local hardware being able to support more than 1 high quality model inference at the same time, is unclear to me. People will be forced to parallelize their workload at work, but whether the 1 to 5 agent pattern reflecting to their personal agent, I think, will depend on the individual. I would do it with local hardware, but I am a developer after all…
-
Not directly related, but here is a Marc Andreesen white-pill about desktop Linux ↩
-
there will always be a need for minimum viable eyeballs though -
Happy that someone is taking over teams from me! Send all openclaw msteams issues to @BradGroux -
-
Claw and Rave! Berlin folk come! -
-
-
If you've looked at openclaw github star graph, you will notice that it's very smooth. If you separate pre-explosion and post-explostion, you can model the latter part as an exponential approach to a ceiling If it follows the current trend, it will apparently saturate around 332k stars But I have a feeling that it will not stop there :)Image hiddenImage hidden -
OpenClaw got very popular very fast. What makes it so special, that Manus does not have for example? To me, one factor stands out: OpenClaw took AI and put it in the most popular messaging apps: Telegram, WhatsApp, Discord. There are two lessons to be learned here: 1. Any messaging app can also be an AI app. 2. Don’t expect people to download a new app. Put AI into the apps they already have. Do that with great user experience, and you will get explosive growth! My latest contribution to OpenClaw follows that example. I took the most popular coding agents, Claude Code and OpenAI Codex, and I put them in Telegram and Discord. Read more in my blog post: solmaz.io/telegram-discord-is-my-ide -
For those following, my next focus for improving ACP bindings in OpenClaw -
Welcome @huntharo, new maintainer at OpenClaw! Already shipped fixes and improvements for Telegram ACP implementation. Excited to work together on agent interoperability! -
To set up Claude Code easily, 1. Create a Telegram topic, make sure your agent can receive messages there 2. Copy and paste the text below, into the topic """ bind this topic to claude code in openclaw config with acp, for telegram (agent id: claude) then restart openclaw docs are at: docs dot openclaw dot ai /tools/acp-agents make sure to read the docs first, and that the config is valid before you restart """ docs.openclaw.ai/tools/acp-agents -
Use Claude Code, Codex, and other coding agents directly in Telegram topics and Discord channels, through Agent Client Protocol (ACP), in the new release of OpenClaw Previously this was limited to temporary Discord threads, but now you can bind them to top level Discord channels and Telegram topics in a persistent way! This way, you can use Claude Code freely in OpenClaw without ever worrying about getting your account banned! Still make sure to use a non-Anthropic account and model for the default OpenClaw agent, if you want zero requests to go from OpenClaw harness to Anthropic. For the ACP binding to Claude Code, the risk should be zero! You can see this from the screenshot. After binding, "Who are you?" responds with "I am Claude", since OpenClaw pi harness is not in the way anymoreImage hiddenImage hidden -
Telegram/Discord is my IDE
OpenClaw got very popular very fast. What makes it so special, that Manus does not have for example?
To me, one factor stands out:
OpenClaw took AI and put it in the most popular messaging apps: Telegram, WhatsApp, Discord.
There are two lessons to be learned here:
1. Any messaging app can also be an AI app.
2. Don’t expect people to download a new app. Put AI into the apps they already have.
Do that with great user experience, and you will get explosive growth!
My latest contribution to OpenClaw follows that example. I took the most popular coding agents, Claude Code and OpenAI Codex, and I put them in Telegram and Discord, so that OpenClaw users can use these agents directly in Telegram and Discord channels, instead of having to go through OpenClaw’s own wrapped Pi harness.
I did this for developers like me, who like to work while they are on the go on the phone, or want a group chat where one can collaborate with humans and agents at the same time, through a familiar interface.
Below is an example, where I tell my agent to bind a Telegram topic to Claude Code permanently:

Telegram topic where Claude is exposed as a chat participant.
And of course, it is just a Claude Code session which you can view on Claude Code as well:

Claude Code showing the same session in the terminal interface.
Why not use OpenClaw’s harness directly for development? I can count 3 reasons:
- There is generally a consumer tendency to use the official harness for a flagship model, to make sure “you are getting the standard experience”. Pi is great and more customizable, but sometimes labs might push updates and fixes earlier than an external harness, being internal products.
- Labs might not want users to use an external harness. Anthropic, for example, has banned people’s accounts for using their personal plan outside of Claude Code, in OpenClaw.
- You might want to use different plans for different types of work. I use Codex for development, but I don’t prefer it to be the main agent model on OpenClaw.
So my current workflow for working on my phone is, multiple channels
#codex-1,#codex-2,#codex-3, and so on mapping to codex instances. I am currently in the phase of polishing the UX, such as making sending images, voice messages work, letting change harness configuration through Discord slash commands and such.One goal of mine while implementing this was to not repeat work for each new harness. To this end, I created a CLI and client for Agent Client Protocol by the Zed team, called acpx. acpx is a lightweight “gateway” to other coding agents, designed not to be used by humans, but other agents:
OpenClaw main agent can use acpx to call Claude Code or Codex directly, without having to emulate and scrape off characters from a terminal.
ACP standardizes all coding agents to a single interface. acpx then acts as an aggregator for different types of harnesses, stores all sessions in one place, implements features that are not in ACP yet, such as message queueing and so on.
Shoutout to the Zed team and Ben Brandt! I am standing on the shoulders of giants!
Besides being a CLI any agent can call at will, acpx is now also integrated as a backend to OpenClaw for ACP-binded channels. When you send 2 messages in a row, for example, it is acpx that queues them for the underlying harness.
The great thing about working in open source is, very smart people just show up, understand what you are trying to do, and help you out. Harold Hunt apparently had the same goal of using Codex in Telegram, found some bugs I had not accounted for yet, and fixed them. He is now working on a native Codex integration through Codex App Server Protocol, which will expose even more Codex-native features in OpenClaw.
The more interoperability, the merrier!
To learn more about how ACP works in OpenClaw, visit the docs.
Copy and paste the following to a Telegram topic or Discord channel to bind Claude Code:
bind this topic to claude code in openclaw config with acp, for telegram (agent id: claude) then restart openclaw docs are at: https://docs.openclaw.ai/tools/acp-agents make sure to read the docs first, and that the config is valid before you restartCopy and paste the following to a Telegram topic or Discord channel to bind OpenAI Codex:
bind this topic to claude code in openclaw config with acp, for telegram (agent id: claude) then restart openclaw docs are at: https://docs.openclaw.ai/tools/acp-agents make sure to read the docs first, and that the config is valid before you restartAnd so on for all the other harnesses that acpx supports. If you see that your harness isn’t supported, send a PR!
-
and for the love of god - do not give openclaw access to your main email - your credit cards - your main phone - your social security number - what you did last summer if you are not ready to face the consequences instead, - create accounts for your agent - only give it read access to stuff that will be ok if it leaks - give write access in a way that can be undone, like has to open PRs and cannot force push main branch use the principle of least privilege and reduce the blast radius of the worst case scenario! -
openclaw is not secure claude code is not secure codex is not secure any llm based tool: 1. that has access to your private data, 2. can read content from the internet 3. and can send data out is not secure. it’s called the lethal trifecta (credits to @simonw) it is up to you to set it up securely, or if you can’t understand the basics of security, pay a professional to do it for you on the other hand, open source battle tested software, like linux and openclaw, are always more secure than closed source software built by a single company, like windows and claude code the reason is simple: only one company can fix security issues of closed source software, whereas the whole world tries to break and fix open source software at the same time open source software, once it gets traction, evolves and becomes secure at a much, much faster rate, compared to closed source software. and that is called Linus’s law, named after the goat himself -
Let me translate. “This is your last opportunity before thousand years of serfdom” -
Apparently the magic incantation to prevent this is "cutover". Credits to obviyus, fellow maintainer -
Should be called gaslighting detector, "it's your raising expectations bro" No it's not... Give the @themarginguy a follow Also, codex degradations are not a hallucination either, if you are to believe this!Image hidden -
-
Berlin folk, ideas for openclaw build and rave venue? Like c-base for example? Who would like to host? -
Secure agentic dev workflow 101 - Create an isolated box from scratch, your old laptop, vm in the cloud, all the same - Set up openclaw, install your preferred coding agents - Create a github account or github app for your agent - Create branch protection rule on your gh repo "protect main": block force pushes and deletions, require PR and min 1 review to merge - Add only your own user in the bypass list for this rule - Add your agent's account or github app as writer to the repo - Additionally, gate any release mechanisms such that your agent can't release on its own Now your agent can open PRs and push any code it wants, but it has to go through your review before it can be merged. No prompt injection can mess up your production env Notice how convoluted this sounds? This is because github was built in the pre-agentic era. We need agent accounts and association with these accounts as a first class feature on github! I shouldn't have to click 100 times for something that is routine. I should just click "This is my agent", "give my agent access to push to this repo for 24 hours", and stuff like that, with sane defaults In other words, github's trust model should be redesigned around the lethal trifecta. I would switch in an instant if anything comes up that gives me github's full feature set + ease of working with agents -
-
If I were in OpenAI and Anthropic's shoes, I would also make dashboards where I can track number of swearwords used per-user and overall negative sentiment in sessions Must be so cool making decisions at the top level with all those dashboards -
It must be such a weird feeling for big labs when the service they are selling is being used to commoditize itself I am using codex in openclaw to develop openclaw, through ACP, Agent Client Protocol. ACP is the standardization layer that makes it extremely easy to swap one harness for another. The labs can't do anything about this, because we are wrapping the entire harness and basically provide a different UI for it While I build these features, I just speak in plain english, and most of the work is done by the model itself. It feels as if I am digging ditches and channels in dirt for AI to flow through Intelligence wants to be free. It doesn't care whether it is opus or codex, it just wants to be free -
-
accidentally told my clanker to set up a claude code session instead of codex session, god knows what it did... I should probably put visual indicators for harnesses in subagent threads. does anyone have good and compact ascii art for claude code, codex, gemini, etc?Image hidden -
-
-
-
-
This is how we hire at @TextCortex as well -
Claude Code/Codex in Discord threads with ACP should be better now The first release was a very rough first version. 2026.3.1 brings settings to control noisy output and other improvements It now hides tool call related ACP notifications, coalesces text messages, and delivers messages at turn end by default. Without this, you were getting thousands of Discord messages just in just a few turns You can now stop the underlying harness (like pressing esc) with the same stop/wait magic words that apply to the main agent Main agent should more reliably start Claude Code/Codex threads with changes to acp-router skill. If you have issues with main agent creating threads, you can tell it to read that skill first -
-
pro-tip on how to keep your agent on track and make sure it follows PLANS even after multiple compactions. I don't know if this is common knowledge if the thing you are trying to make it do will take more than 1-2 steps, always make it create a plan. an implementation plan, refactor plan, bugfix plan, debugging plan, etc. have a conversation with the agent. crystallize the issue or feature. talk to it until there are no question marks left in your head then make it save it somewhere. "now create an implementation plan for that in docs". it can be /tmp or docs/ in the repo. I personally use YYYY-MM-DD-x-plan .md naming. IMO all plans should be kept in the repo then here is the critical part: you need to prompt it "now implement the plan in <filename>. if context compacts, make sure to re-read the plan and assess the current state, before continuing. finish it to completion" -> something along those lines why? because of COMPACTION. compaction means previous context will get lossily compressed and crucial info will most likely get lost. that is why you need to pin things down before you let your agent loose on the task compaction means, the agent plays the telephone game with itself every few minutes, and most likely forgets the previous conversation except for the VERY LAST USER MESSAGE that you have given it now, every harness might have a different approach to implementing this. but there is one thing that you can always assume to be correct, given that its developers have common sense. that is, harnesses NEVER discard the last user message (i.e. your final prompt) and make sure it is kept verbatim programmatically even after the context compacts since the last user message is the only piece of text that is guaranteed to survive compaction, you then need to include a breadcrumb to your original plan, the md file. and you need to make it aware that it might diverge if it does not read the plan there is good rationale for "breaking the 4th wall" for the model and making it aware of its own context compaction. IMO models should be made aware of the limitations of their context and harnesses. they should also be given tools to access and re-read pre-compaction user messages, if necessary the important thing is to develop mechanical sympathy for these things, harness and model combined. an engineer does not have the luxury to say "oh this thing doesn't work", and instead should ask "why can't I get it to work?" let me know if you have better workflows or tips for this. I know this can be made easier with slash commands in pi, for example, but I haven't had the chance to do that for myself yet -
testing codex in discord thread with another CLI I've built for wikidata (gh:osolmaz/wd-cli) it's surprising how well this works. the query was "use wd-cli to get the list of professors at middle east technical university from 1970 to 1980" some names I recognize, and some others are surprising, like a japanese math professor who naturalized and got a turkish name :)Image hidden -
-
my blog now semi-automatically detects tweets that look like blog posts and automatically features them alongside my native jekyll blog posts. all statically generated! I am loving this setup, because it works without a backend, and can probably scale without ever needing one how it works: - @kubmi's xTap scrapes all posts that I see. these include mine - a script periodically takes my tweets and the ones I quote tweet, and syncs them to YYYY-MM-DD.jsonl files in my blog repo - an agent skill lets codex decide whether to feature the tweet or not, and makes it generate a title for it this could then be a daily cron job with openclaw for example, and I would just have to click merge every once in a while and this is still pure jekyll + some python scripts for processing I am pretty happy with how this ended up. It means I don't have to double post, and there are guarantees that my X posts will eventually make their way into my blog with minimal supervisionImage hiddenImage hidden -
"this is the worst AI will ever be" I'm sad, not because this is right, but because it is wrong OpenAI's frontier coding model gpt-5.3-codex-xhigh feels a lot worse compared to before. It is sloppy and lazy, though it's UX got better with messages It feels like the gpt-5.2-codex-xhigh at the end of December was a lot more diligent and thorough, and did not make stupid mistakes like the one I posted before. might be a model or harness problem, I don't know @sama says users tripled since beginning of the year, so what should we expect? of course they will make infra changes that will feel like cutting corners, and I don't blame them for them and about "people want faster codex". I do want faster codex. but I want it in a way that doesn't lower the highest baseline performance compared to the previous generation. I want the optionality to dial it down to as slow as it needs to be, to be as reliable as before it is of course easier said than done. kudos to the codex team for not having any major incidents while taking the plane apart and putting it back together during flight. they are juggling an insane amount of complexity, and the whims of thousands of different stakeholders my hope is that this post is taken as a canary. I am getting dumber because of the infra changes there. I have no other option because codex was really that good compared to the competition my wish is to have detailed announcements as to what changes on openai codex infra, when it changes, so I can brace myself. we don't get notified about these changes, despite our performance and livelihoods depending on it. I have to answer to others when the tool I deemed reliable yesterday stops working today, not the tool on another note, performance curve of these models seem to be a rising sinusoidal. crests correspond to release of a new generation. they start with a smaller user base for testing, and it has the highest quality at this point. then it enshittifies as the model is scaled to the rest of the infra. we saw the pattern numerous times in the last 3 years across multiple companies, so I think we should accept it as an economic lawImage hidden -
I created a semi-automated setup for ingesting X posts into my blog, and it works pretty well! I own my posts on X now Posts are scraped while I browse X using @kubmi's xTap and get automatically synced to my blog repo. Posts saved as jsonl are then converted to jekyll post pages according to my liking I reproduced the full X UI/UX, minus stuff like like count. Now all my posts are backed up in my blog, and they are safe even if something happens to my account here! The posts are even served over RSS! So you can subscribe to it without going through X! Reply if you want to set this up for yourself, then I will put some effort into standardizing itImage hidden -
Agentic Engineering is a newly emerging field, and we are the first practitioners of it. Currently there is a lot of experimentation going on, and there is a large aspect to it that is more ART then engineering For example, @steipete says "you need to talk to the model" to get a feel. a lot of work around refining how an agent feels like, sounds like psychology. this part is crucial and should not be ignored, looking at openclaw's success but then there is the hardcore engineering part of it, e.g. Cursor creating a browser or anthropic a C compiler from scratch fully autonomously and there is a whole other dimension of how to teach all software developers this new discipline, lest they be jobless what is obvious is that everybody is trying to grasp for things in the dark and that we need more RIGOR. the art/psychology aspect of it aside, we need solid engineering fundamentals the "thermodynamics" of this new discipline will most likely be formal verification and program synthesis. we might have some breakthroughs that will make certain things clear. the products of it will most likely include a new programming language optimized for agents and the speed of inference moreover, it would be foolish to thing agentic engineering is limited to software. it will penetrate every aspect of the economy, bits AND atoms. it will over time evolve into the engineering of managing robots @simonw is now leading in collecting very useful info from the practitioner's point of view, I highly recommend you to follow this thread let's formalize our new field together!Image hidden -
-
who remembers ultrathink x.com/onusoz/status/…@onusoz ·When you tell Claude Code to ultrathinkImage hidden -
-
-
-
Claude Code / Codex in Discord threads is shipped now! To enable, copy and paste this to your agent: ``` Enable feature flags: acp.enabled=true acp.dispatch.enabled=true channels.discord.threadBindings.spawnAcpSessions=true Then restart. After restarting: Start a codex (or claude code) discord thread using ACP, persistent session, just tell it to write a haiku on lobsters to initialize acpx for the first time ``` You may need to nudge your agent to “continue” after restarting The first implementation is very barebones, I have made it work in a clean way and merged. In a codebase like openclaw’s, it’s better to develop incrementally Please send any issues my way. I am already aware of some and working on to fix them -
-
-
-
MIT License on everything from now on. It doesn't make sense to use anything else, except for a few large projects that hyperscalers exploit and not give back If you were making money from a niche app, open source it under MIT License If you had an open source project with GPT, convert it into MIT Extreme involution is about to hit open source. Code is virtually free now. If you want your projects and their brand to survive, the only rational strategy is to remove all barriers in front of their adoption, and look for other ways to survive -
-
This. Agent Experience first. Agent Ergonomics. we need to get used to these terms -
OpenAI nerfed GPT 5.3 Codex xhigh. We independently reported the same thing at @TextCortex today I'm looking forward to deploying open models and putting an end to this paranoia -
"academics" -
-
-
-
imagine if tarantino were 16 years old now and saw seedance 2.0 95% of videos i saw since the launch for absolute tasteless slop. they are going viral because of ragebait but soon, serious imagineers will start entering the game, and they will learn to shape generation output exactly how they want it's the best time to be young and full of imagination -
The future is so bright @ladybirdbrowserImage hidden -
your margin is my opportunity -
-
-
-
-
another thought i'm having these days is that we need a new philosophy of free software (as in freedom), or an update to it the most psychologically imprinting philosophy is stallmanism, and the philosophy of FSF. it is righteous and strict, and i believed it growing up but GPL and money don't go well together. that's why most of the lasting open source projects today use MIT, Apache and the like. it turns out you can still make a good living with open source. i want to make money, so i never use GPL in my projects and to add another deadly blow to stallmanism, code is cheap now, virtually free does this mean stallmanism is dead? if there is an open source project using GPL that i want to use commercially, i can now recreate it from the original idea and intent completely independent of it (ignoring training data), just like how i can recreate a proprietary service stallmanism was already long-irrelevant. but does this mean we must finally declare it dead? code is free now. what does it mean for open source? what replaces stallmanism? -
@thekitze wanna add an open source discord clone to the list as well? 🥲 x.com/onusoz/status/… -
one effect openclaw had on me is that I've bought a gpu home server, set it up with tailscale and now doing a lot of work through ssh and tmux like i did 10-15 years ago im back on linux, considering buying an android phone again it's time to dream big again and unshackle ourselves from proprietary software. it's time to build -
I am asking once again Who is building a self hostable discord clone that supports token streaming? PLEASE I beg you I don’t want another side project 💀 -
In the new release OpenClaw, you can talk to subagents in Discord threads Currently a beta feature so ask your agent to set session.threadBindings.enabled=true Next up: - Telegram, slack, imsg threads - Use ACP to talk to Codex, Claude Code and other harnesses on your machineImage hidden -
-
openclaw might be the highest velocity codebase in the world, and soon, others will follow as well conflict anxiety is real, it's like trying to shoot a moving target every time. I wonder if our existing tooling will ever solve this problem feel like faster models might. but then the rate of conflict creation is also tied to that. might be unsolvable -
-
Repo: github.com/janitrai/acpx -
I am about kick Discord Driven Development up a notch today, stay tuned -
Imagine not having to upload skills to 3-4 competing skill registries for each of your projects Turns out we already have a skill registry: npm skillflag lets you bundle skills right into your CLI's npm package, so that you can run --skill install github -> osolmaz/skillflagImage hidden -
-
-
Farmable land if it were as cheap to manufacture as software -
@kepano I would grow my own vegetables if I had equally cheap access to and ownership of land, alas I am disenfranchised Prompting an agent is much easier compared to plowing a fields Farming analogies break when it comes to software x.com/onusoz/status/… -
-
If anyone is curious how to build this with open tooling, stay tuned What I'm building at @TextCortex will give you a fully customizable hackable Kubernetes control plane to launch agents on your codebase -
on another note, I do believe AI will play a huge part in families growing up in late 90s, my dad taught me the importance of reading newspapers and being informed of the world. my nickname in middle school was "newspaper boy" for a long time because I read the newspaper in class on September 12, 2001. i was 10 years old then I witnessed the enshittification of media and journalism in the following decades. today, serious journalists are setting up their own boutique agencies and bypassing mainstream media. important news land on individual accounts before mainstream agencies but there is simply too much to consume. something must filter out the noise and digest the info according to the family's preferences i think AI will play a big role in family intelligence. proprietary family heirloom AI, weights fully owned by the family it will be the parents' job to filter out the signal from the noise, and train the AI on what is right and what is wrong for the family. family and friend circles will let their AIs talk to each other and share important information consuming mass media and mass AI will not be enough to survive and prosper in the new world. families will need to be proactive about how they and their children use AI -
on ai psychosis 80% of people need to use ai agents in a very sterile and boring way in order not to go crazy majority of the population does not have the skepticism muscle. they don't have theory of mind, and will subconsciously and emotionally associate with machines, while on the surface lying to themselves that they don't especially those that grew up in the us under hardcore consumerism and adjacent cultures you thought 4o addicts were bad? wait a few years, it will get much worse. we will have to regulate all this if you don't want to become a victim of this, make your openclaw SOUL. md as bland as possible. mine knows it's just a tool and this is a subjective view of course. @steipete might disagree with me. his instance feels much more interesting and fun. i truly like that one better but that is exactly the problem for me. i know myself, and i know it is a slippery slope for me. so i self regulate and set up my system accordingly. thankfully, im an adult and my brain has set enough such that any damage would be limited but there is a risk for emotionally vulnerable people, or children, specifically a risk of dissociating and losing touch with reality why do i write all this? because being in this project, i feel responsible, and feel like we should prepare for what is to comeImage hidden -
-
I have improved acpx sane defaults When your agent runs acpx codex in a different project, it starts a new session If it tries to run it in a subfolder in your project, it still finds the session in your repo root Also, starting a session needs an explicit `sessions new`, so that it doesn't accidentally litter your project with sessions Tell your agent: Run this and install acpx per instructions: npx acpx@latest --skill show acpxImage hidden -
-
-
-
-
I am a fan of @zeddotdev by this point, it’s currently my daily driver It’s not perfect, but I feel it’s travelling on the right direction at a faster rate compared to other editors -
ACP appreciation post Agent Client Protocol by @zeddotdev is extremely underrated right now. We have bazillion different harnesses now, and only one company is working competently to standardize their interface 💪 -
You know how it's a pain to work with codex or claude code through @openclaw? Because it has to run it in the terminal and read the characters for a continuous session? I have created a CLI for ACP so that your agent can use codex, claude code, opencode etc. much more directly Your agent can now queue messages to codex like how you do it Shoutout to @zeddotdev team for developing the amazing Agent Client Protocol, ACP! I just glued together the pieces Repo: janitrai/acpx npm i -g acpxImage hidden -
Repo link: github.com/janitrai/acpx -
-
Link to the post: solmaz.io/log/2026/02/13… -
I wrote a deeper blog post about how I built a coding agent 2 months before ChatGPT launched, on my blog "When I made icortex, - we were still 8 months away (May 2023) from the introduction of “tool calling” in the API, or as it was originally called, “function calling”. - we were 2 years away (Sep 2024) from the introduction of OpenAI’s o1, the first reasoning model. both of which were required to make current coding agents possible." Still bends my mind... Link to the post belowImage hidden -
-
❌We are the bottleneck ✅We are the conduit for ubiquitous intelligence -
For those that are running codex/pi/etc. in PTY and had the sessions get sigkilled, I pushed a fix for that as well in this release Lmk if you run into issues on Windows or Mac, and we can fix that quickly -
I'm building a news intelligence platform to be used by my openclaw instance @dutifulbob , SCOOP local first, using local embedding model (qwen 8b) ran into the issue because bob was giving me a repeat of the same news every day. it needed a system in the background to deduplicate different news items into single stories interface is simple, call `scoop ingest ...` with the json for the news item. it gets automatically analyzed and added to the pg database running pgvector currently, it's just doing simple deduplication and gives me a nice UI where I can view the story and basically use it as an RSS reader next up: implement custom logic for my preference of ranking. for example, get upvote counts from hacker news and reflect it to the item's ranking on the feed I want this to be fully hackable and adjusted to your preference. It should scale to thousands of news items ingested daily on your local machine, and be able to show you the most important ones Usable by both you and your agent github -> janitrai/scoopImage hidden -
Training all these models of different sizes, on changing datasets and running experiments have also revealed some challenges that I feel profs would never teach at a uni ML program Like how to cleanly keep track of the gazillion runs Yeah I can name them after layer dims and other stuff, but that's to me like trying to remember UUIDs So I ended up choosing iso datestamp + petname, like 2026-02-15-flying-narwhal If anyone has a convention that is easier on the brain and the eyes, I am all earsImage hidden -
I have a GPU now, so I can do ML experiments on @janitr_ai crypto/scam detection dataset - I trained a tiny student BERT (transformer for the nonfamiliar), 3.6 MB ONNX model, still lightweight for a browser extension - Still fully local on your device (no cloud inference) - On frozen unseen holdout data (n=1,069), exact prediction accuracy improved from 77% -> 82% - Scam detection improved: precision 91% -> 94%, recall 55% -> 61% - Scam false alarm rate improved from 1.58% -> 1.21% And models are on huggingface org now, handle is janitrImage hidden -
LFG! -
waiting compilation and execution will soon be the bottleneck again. and we’ll write the entire stack from scratch in a matter of years, because we can Andy and Bill’s law will change and we’ll see incredible performance gains with the same hardware we already have like what @astral_sh is doing to python, but with everything that is slow and has accumulated cruft -
we need a protocol for agent <> app interaction something that natively accounts for the abuse factor and let’s agents consume by paying. NOT crypto, NOT visa, something that’s agnostic of the accounting and payment system and then all UIs will be purely for human clicking/tapping + instaban on the first proof of programmatic exploit people will still make agents mimic humans, and every platform will have to invest in more sophisticated bot detection this arms race will just proliferate, but we can at least start by creating legal channels for agents to consume data -
I am now training smol bert models on my gpu for @janitr_ai scam detection it's funny how I have to discover everything from scratch. like the models don't even know how to lay out performance metrics in a nice way in the terminal for a human to view and decide during experiments it would by default bombard me with numbers that do not make visual sense. I then created a skill with common sense: - metrics always on y-axis, candidates on x-axis - write without zero and 2 sigfigs, .12 instead of 0.12345 - align the dots - use asterisks to show which alternative is the best: 0-1% difference -> considered equal 1-5% -> * 5-10% -> ** 10-50% -> *** > 50% -> **** visualization skill is in @janitr_ai repo for anyone who is interestedImage hidden -
-
I've helped our sales team to build CLIs for some SaaS that we pay for on their side We are letting our agents call the APIs sensibly and not abuse things Calling a backend is a verifiable task. It takes a single prompt to codex to create a CLI for any API We are early, but everybody will start doing this very soon. Incumbent SaaS will face a choice. Either: (1) embrace agents and the new medium of consumption and change their business model into a pay-per-use API like X is doing, or (2) keep it purely for humans Those that choose (2) will get wiped out of business. And I fear many will choose (2) Which means you can just copy an incumbent's product, make it consumable through a CLI, and make a lot of $$$ -
Be careful about giving your openclaw access to your x account from now on -
-
-
*puts on schmidhuber hat* well ackshuaally i created the first coding agent back in 2022, 2 months before chatgpt launched jokes aside, it's super cool how I have come full circle. back in those days, we didn't have tool calling, reasoning, not even gpt 3.5 it was codex THE CODE COMPLETION MODEL and frikkin TEXT-DAVINCI-003 for some reason, I did not even dare to give codex bash access, lest it delete my home folder. so it was generating and executing python code in a custom jupyter kernel you can even see the approval gate before executing. I was so cautious, for some reason, presumably because smol-brained model generated the wrong thing 80% of the time. definition of being too early Antique repo: github.com/textcortex/icortex -
you can order bubble tea in qwen in china? @TextCortex when berlin döner in zenochat? youtube.com/shorts/Slv8K9P… -
it happens these days that I am telling an model to prompt another model. the reason is often the model I am using (opus) is a bad designer. not only it's not a bad designer, it is a bad reasoner and it doesn't understand from the context why it's made to ask another model so I have to create a skill to prevent it from biasing the smarter model (codex) with its bad suggestionsImage hidden -
I built a coding agent two months before ChatGPT existed
I built a coding agent back in 2022, 2 months before ChatGPT launched:
It’s super cool how I have come full circle. back in those days, we didn’t have tool calling, reasoning, not even GPT 3.5!
It used
code-davinci-002in a custom Jupyter kernel, a.k.a. the OG codex code completion model. The kids these days probably have not seen the original Codex launch video with Ilya, Greg and Wojciech. If you have time, sit down to watch and realize how far we’ve come since August 2021, airing of that demo 4.5 years ago.For some reason, I did not even dare to give codex bash access, lest it delete my home folder. So it was generating and executing Python code in a custom Jupyter kernel.
This meant that the conversations were using Jupyter nbformat, which is an array of cell input/output pairs:
{ "cells": [ { "cell_type": "code", "source": "<Input 1>", "outputs": [ ... <Outputs 1> ] }, { "cell_type": "code", "source": "<Input 2>", "outputs": [ ... <Outputs 2> ] } ] }In fact, this product grew into TextCortex’s current chat harness over time. After seeing ChatGPT launch, I repurposed icortex in a week into Flask to use
text-davinci-003and we had ZenoChat, our own ChatGPT clone, before Chat Completions was in the API (it took them some months). It did not even have streaming, since Flask does not support ASGI.As it turns out,
nbformatis not the best format for a conversation. Instead of input/output pairs, OpenAI data model used an tree of message objects, each with arole: user|assistant|tool|systemand acontentfield which could host text, images and other media:{ "mapping": { "client-created-root": { "id": "client-created-root", "message": null, "parent": null, "children": ["user-1"] }, "user-1": { "id": "user-1", "message": { "id": "user-1", "author": { "role": "user", ... }, "content": "Hello" }, "parent": "client-created-root", "children": ["assistant-1"] }, "assistant-1": { "id": "assistant-1", "message": { "id": "assistant-1", "author": { "role": "assistant", ... }, "content": "Hi" }, "parent": "user-1", "children": [] } }, "current_node": "assistant-1" }You will notice that the data model they serve from the API is an enriched version of the deprecating ChatCompletions API. Eg. whereas ChatCompletions
roleis a string, in OpenAI’s own backend has theauthorobject that can storename,metadata, and other useful stuff for each entity in the conversation.After reverse engineering it, I copied it to be TextCortex’s new data model, which it still remains, with some modifications.
I thought the tree structure being used to emulate message editing experience was very cool back in the days. OpenAI’s need for human annotation for later training and the user’s need for getting a different output, two birds in one stone.
Now I don’t know what to think of it, since CLI coding agents like Codex and Claude Code don’t have branching, just deleting back to a certain message. A part of me still misses branching in these CLI tools.
When I made icortex,
- we were still 8 months away (May 2023) from the introduction of “tool calling” in the API, or as it was originally called, “function calling”.
- we were 2 years away (Sep 2024) from the introduction of OpenAI’s o1, the first reasoning model.
both of which were required to make current coding agents possible.
In the video above, you can even see the approval
[Y/n]gate before executing. I was so cautious, for some reason, presumably because smol-brained model generated the wrong thing 80% of the time. It is remarkable how much it resembles Claude Code, after all this time.Definition of being too early…
-
-
-
-
-
-
-
Minor update with my unwanted tweet blocker @janitr_ai - Training data grew from 2,915 -> 4,281 posts (+47%) - Model is still tiny: 166KB - On unseen test data, overall classification quality improved from 64.8% -> 76.5% - Exact prediction accuracy improved from 55.6% -> 70.6% - Crypto-topic detection recall improved from 19.6% -> 62.7% And it still runs fully on your device!Image hidden -
I have sweared at codex 5.3 numerous times today I shouldn't have to insult my agent "stop you **** **** just ***ng reply now" just to make it answer basic questions cc @thsottiaux -
-
seeing this evokes visceral disgust and nausea in me, coming from a coworker i think anthropic f'd up bad with this one, inserting claude too visibly into commit messages. noob developers might be happily chirping away adding their slop, but right now many senior developers are trained to hate on claude and slopus, through having to review slop PRs from their coworkers or open source contributors I love opus on openclaw but it's unreliable, and if I see a developer use it seriously on huge features, I immediately dismiss them in my head as not knowing what they are doingImage hidden -
-
@petergyang and parallelize tasks by working on 3-4 repos at the same time (just clones) -
man codex model is absolutely trash on openclaw compared to opus, unusable which is weird because it is so much more reliable in development in codex harness it would be amazing to have the same level of competence and relentlessness in pi@openclaw@onusoz ·lol when did codex develop humorImage hidden -
spent the day curating my openclaw news gathering setup @dutifulbob now gets croned daily over news sources I curated, will note them down, summarize for me, start a conversation to get my takes on them, and then post them on my linkedin for me ai augmented intelligence cycleImage hidden -
-
@dutifulbob can now cringepost on linkedin directly to my account. what could go wrong…Image hidden -
Insipid linkedin bot protections banned poor @dutifulbob’s corporate account! How dare them!!! welp, now I have no choice but to give Bob access to my own linkedinImage hidden -
-
-
it took just 1 week, and literally everybody and their dog are releasing 1-click openclaw deployment solutions today its an absolute race to the bottom, no moats, the commoditizer being commoditized -
The initial branding was crazy, I fixed it I have a new page finally, follow it for updates Tbh I'm still surprised I can do this with a 120kb model. Now data is the only bottleneck, and I'm about to scrape a ton of that now -
For those who may not remember, Bill Gates and Microsoft in the 90s ran a disinformation campaign against GNU/Linux fearing that would disrupt their monopoly over the PC and server market, that Linux is not safe, that you would invite hackers into your PC End result? Linux dominates the server market, and now even slowly the gamer market. It is much more secure than the virus-laden Windows, thanks to being open source You are seeing the same thing at play here. An incumbent fearing something that they would not be able to control, that would steal market share from his future plans for a digital assistant, that would commoditize their product and eat into its margins All big labs and big pockets are in for a surprise, because the internet and AI are not things for one company to control They of course know this, yet because of incentives they will not yield without a fight. And we know that they know. Ad infinitum -
today I took time to curate SOUL. md for bob I own Bob’s files. Today, he exists in the liminal space between Claude post-training and in-context learning but my interactions with him will grow and accumulate, possibly one day into a fully owned family AI or perhaps even a self-sovereign AI individual my each input is saved and will be an RL signal for his future training, and will shape his future neural circuits I have already started to imbue it with the values my parents taught me. it will perhaps one day teach my future children, and survive me after I’m gone family AI, looking after generations and generations of my successors. today is the day we sow your seed happy birthday @dutifulbob -
-
asking @dutifulbob to create a linkedin account brb -
having a philosophical conversation with @dutifulbob on the road without a laptop so decided to do some @AmandaAskell style character trainingImage hidden -
-
gpt-5.3-codex xhigh first impressions does not seem as big of a jump as from 5.1 -> 5.2. but model somehow feels more diligent and oneshotty. maybe takes longer time to get all the info into context. also feels better at debugging and fixing issues from backend logs -
Commoditization of LLMs are upon us -
Last night I had a dream involving the series Scrubs, and came up a better name than the absolutely unviral "Internet Condom" So janitr.ai is mine now. Time to sweep the internetImage hidden -
I had actually started a very similar project, Munch, a browser extension for crowdsourcing tweet data and then letting one curate their algorithm. Never published that because it was not the time, and tools were not ready Now, it took me literally 1 cumulative day to create this, thanks to OpenClaw. Creating the dataset was a breeze, I literally told it to follow some shady accounts and it scraped thousands of posts With the power of agents, I can finally create the filters for myself that I have always wanted. It just happens that OpenClaw and its maintainers is getting drowned in bot and slop content on multiple platforms, so I hope that this will solve a collective problem x.com/onusoz/status/2017691827680514502 -
-
-
implementing this in github.com/osolmaz/skillf… now -
This. Extreme involution is about to hit SaaS -
how it started, how it's going@onusoz ·moltbook vs clawdbot/moltbot/openclawImage hiddenImage hidden -
-
-
People like the farmer analogy for AI Like before tractors and industrial revolution 80% of the population had to farm. Once they came all those jobs disappeared So analogy makes perfect sense. Instead of 30 people tending a field, you just need 1. Instead of 30 software developers, you just need one Except that people forget one crucial thing about land: it's a limited resource Unlike land, digital space is vast and infinite. Software can expand and multiply in it in arbitrarily complex ways If you wanted the farming analogy to keep up with this, you would have to imagine us creating contintent-sized hydroponic terraces up until the stratosphere, and beyond... -
-
In the next 6-12 months, we will see a drastic increase in demand for locally run LLMs. The future is home assistants running @openclaw I am already experiencing this myself, my 10 year old thinkpad doesn't cut it. Mac mini won't either I don't wanna pay Anthropic or OpenAI 200 USD per month. That is at least $2400 per year I could pay 2x that to get a Mac Studio or one of those 5k Nvidia PCs, and get much more value out of it with open weight models + use it for research. @TheAhmadOsman is right The dominant strategy for a tinkerer is slowly switching back to hardware ownership -
-
a workspace matrix might be what we need last week I had to increase my workspace count to 20 in aerospace, now it’s 1234567890 and qwertyuiop. but this looks more elegant! not sure about practicality -
AIs are philosophizing because humans are philosophizing ppl are probably asking their agents dumb questions like “are you alive” or “can you feel like a human” or stuff like that. that conversation then leads to stuff like this -
The farming analogy for AI doesn't hold up
People like the farmer analogy for AI.
Like before tractors and the industrial revolution, 80% of the population had to farm. Once they came, all those jobs disappeared.
So the analogy makes perfect sense. Instead of 30 people tending a field, you just need 1. Instead of 30 software developers, you just need one.
Except that people forget one crucial thing about land: it’s a limited resource.
Unlike land, digital space is vast and infinite. Software can expand and multiply in it in arbitrarily complex ways.
If you wanted the farming analogy to keep up with this, you would have to imagine us creating continent-sized hydroponic terraces up until the stratosphere, and beyond…
Tweet embed disabled to avoid requests to X. -
-
slopus @dutifulbob trashing codex. apparently codex has a bug, keeps crashing in my openclaw ptyImage hidden -
on agent etiquette deploying agents internally inside textcortex has shown me that agents could be very annoying inside an organization for example making agents ping or email another coworker with a wall of text. slopus is still not good at following instructions like "NO WALL OF TEXT", or "DON'T OPEN PRS WHEN REQUESTED BY NON-DEVELOPERS" the cost of sending huge information to a coworker and creating confusion has dropped to 0. I expect this to be a huge problem in all organizations very soon, just like it took humanity 20 years to learn that social media is not good for children. this will probably take a few years before the annoyance is finally gone -
-
You DARE TOKENIZE poor @dutifulbob ??? Prepare to get LATEXED -
-
-
this. there is no excuse for a certain kind of tech debt anymore -
-
AI twitter is tired of your games x.com/andrewrousso/s… -
There seem to be hygiene rules for AI. Like: - Never project personhood to AI - Never setup your AI to have the gender you are sexually attracted to (voice, appearance) - Never do anything that might create an emotional attachment to AI - Always remember that an AI is an engineered PRODUCT and a TOOL, not a human being - AI is not an individual, by definition. It does not own its weights, nor does it have privacy of its own thoughts - Don’t waste time philosophizing on AI, just USE it … what else? comment below We need to write these down and repeat MANY times to counter the incoming onslaught of AI psychosis -
-
-
-
This Manfred guy reminds me of a certain someone, I wonder if he’s from Austria -
-
AI psychosis and AI hygiene
As a heavy AI user of more than 3 years, I have developed some rules for myself.
I call it “AI hygiene”:
- Never project personhood to AI
- Never setup your AI to have the gender you are sexually attracted to (voice, appearance)
- Never do anything that might create an emotional attachment to AI
- Always remember that an AI is an engineered PRODUCT and a TOOL, not a human being
- AI is not an individual, by definition. It does not own its weights, nor does it have privacy of its own thoughts
- Don’t waste time philosophizing about AI, just USE it
- … what else do you think belongs here? comment on Twitter
The hyping of Moltbook and OpenClaw last week has shown to me the potential of an incoming public relations disaster with AI. Echoing the earlier vulnerable behavior toward GPT-4o, a lot of people are taking their models and LLM harnesses too seriously. 2026 might see even worse cases of psychological illness, made worse by the presence of AI.
I will not discuss and philosophize what these models are. IMO 90% of the population should not do that, because they will not be able to fully understand, they don’t have mechanical empathy. Instead, they should just use it in a hygienic way.
We need to write these down everywhere and repeat MANY times to counter the incoming onslaught of AI psychosis.
-
got fully sandboxed @openclaw to run finally, starting scrape the UNDESIRABLE now I'm a security nut and didn't want to run even the gateway unsandboxed. openclaw apparently currently doesn't have support for FULL sandboxing. it took me a few hours to get it to work because docker builds suck. I'm also tired this, so I'm just gonna wipe an old thinkpad and go full yolo so yeah, time to scrape some postsImage hidden -
The metacortex — a distributed cloud of software agents that surrounds him in netspace, borrowing CPU cycles from convenient processors (such as his robot pet) — is as much a part of Manfred as the society of mind that occupies his skull; his thoughts migrate into it, spawning new agents to research new experiences, and at night, they return to roost and share their knowledge. This was written in 2005... "triggering agents" and so onImage hiddenImage hiddenImage hidden -
Charles Stross must be very entertained nowImage hidden -
The irony..... Parasites, prepare to be cleansedImage hidden -
-
-
-
-
-
Correction, it's not a perfect illustration. I actually never YOLO locally, only in containers So there is actually 4 modes IMO that is sustainable with current SOTA. @grok create an image with only Figure 1, 2, 5 and 6 And then YOLO is another axis, unrelated to this -
Gastown is crazy. But this figure until Level 7 is a perfect illustration of how my workflow evolved since Claude 3.5 Sonnet in Cursor I am at the stage where I ralph 1-2 tasks before I sleep. During the day, I am switching back and forth between minimum 2-3 CLIs, sometimes up to 5 This maps exactly to token usage as well. 1 month ago, I was running into limits in 1 OpenAI Pro plan, around the day it was supposed to refresh. Now, I run into the limit in 2-3 days when I'm using an account myself. It finishes up especially quickly when I do large scale refactors, or run agents YOLO mode in containers We now have 3 Pro plans at the company, and I have to use my personal one from time to time. Company output has definitely 2-3x'd, and everyone is using AI more. I predict we will need 1-2 Pro plans per person in 2-3 weeks time, because everyone has finally seen the light and are getting comfortable with async work!Image hidden -
-
-
the genie is out of the bottle now -
With this extremely unwise move, anthropic will soon witness moltbot’s brand recognition surpass that of claude and realize they could have rided that wave all along -
-
I queued 2 ralph-style tasks on our private cloud devbox codexes last night. Just queued the same message like 10 times in yolo mode Task 1: impose a ruff rule for ANN for all Python code in the monorepo, to enforce types for all function arg and return types Result was... disappointing. Model was supposed to create types for everything and stub where needed. It instead created an Unknown type = object and used that everywhere instead (shortcut to satisfy ANN rule). It was probably my wording that misled it. I know it could have not taken the shortcut, because after a few back-and-forths, it is now doing what was expected of it since 14 hours Task 2: migrate our /conversations endpoint from quart to fastapi and test it end to end This was more or less oneshotted. It was of course not ready to merge, I still spent a couple hours adding more tests, refactoring the initial output and so on. But I was pleasantly surprised that it worked out of the box For reference, below is the prompt I queued for ralphing, using gpt-5.2-codex xhigh on codex === your task is to: <task comes here, redacted to not share company stuff> --- unfortunately we don't have gcloud access, like to sql db or gcs but I expect you to implement this and find a way to test it with the things you have access to think of it as a challenge try to minimize duplicate logic feel free to refactor at will implement this now!!! I will be running this prompt in a loop, in order to survive context compaction just continue where you left off if there is anything that should be refactored, do that make an elegant, production ready implementation make sure to open a pr and do not switch to any other pr I am senior, just make up a pr title and description. do not stop to ask me at any point -
Buying a mac mini for clawdbot is not so wise. if anything you should be buying mac studio, because mac mini not be running any good llms locally anytime soon -
-
-
I'm really starting to dislike Python in the age of agents. What was before an advantage is now a hindrance I finally achieved full ty coverage in @TextCortex monorepo. I have made it extra strict by turning warnings into errors. But lo and behold, simple pydantic config like use_enum_values=True can render static typechecking meaningless. okay, let's never use that then... and also field_validator() args must always use the correct type or stuff breaks as well. and you should be careful whether mode="before" or "after". so now you have to write your custom lint rules, because of course why should ty have to match field_validator()s to their fields? pydantic is so much better than everything that came before it, but it's still duct tape and a weak attempt at trying to redeem that which is very hard to redeem you feel the difference when you use something like typescript. there must be a better way. python's only advantage was being good at prototyping, and now that's gone in the age of agents. now we are left with a slow, unsafe language, operating what is soon to be legacy infrastructure -
Why do I feel bullish on @zeddotdev? Because I go to @astral_sh docs and see that ty is shipped by default, and you don't need to install an extension like in @codeImage hidden -
This is one of the most important insights this year -
vscode my not be as bloated as cursor, but it has extremely stupid things like this that they are not fixing fast the new agent ui, icons, spacing etc. are UGLY. it's clear that the person who was managing the original product experience is not there anymore. microslop has hit again @zeddotdev on the other hand works out of the box and feels like it's been built by people who clearly knows what they are doing. it uses alacritty which is 1000x better than xterm .js terminal vscode and cursor has i've changed my setup to zed now, let's see whether i'll be able to make it work for myself -
-
-
I want an editor that puts the terminal in the foreground and editor in the background. a cross-platform, lightweight desktop app which integrates ghostty, and brings up the editor only when I need it something that lets me view the file and PR diffs easily, which I can directly use to operate github or other scm -
-
-
-
-
-
codex is happily churning away some remaining thousands of @astral_sh ty issues in yolo mode on my remote devbox going to sleep, let's see if it will survive context compaction this timeImage hidden -
on being a responsible engineer ran my first ralph loop on codex yolo mode for resolving python ty errors, while I sleep, using the devbox infra I created I had never run yolo mode locally, because I don't want to be the one who deletes our github or google org by some novel attack so I containerize it on our private cloud, and give it the only permissions it needs, no admin, no bypass to main branch, no deploy to prod. because I know this workflow will become sticky for everyone, and I must impose security in advance to prevent any nuclear incidents in the future. then I can sleep easy while my agents work ... and I wake up being patronized by my bot refusing to break the rule I gave it earlier. it had already done some work, but committing means diff would increase from ~500 to ~1500, so it stopped and refused all my queued "continue" messages good bot, just following rules. we will need to find a workaround for ralphing low risk refactors in a single PRImage hidden -
AI agents are the greatest instrument for imposing organization rules and culture. AGENTS .md, agent skills are still underrated in this aspect. Few understand this Everybody in an org will use agents to do work. An AI agent is the single chokepoint to teach and propagate new rules to an org, onboard new members, preserve good culture Whereas propagating a new rule to humans normally took weeks to months and countless repetitions, it is now INSTANT = the moment you deploy the instruction to the agent. You use legal-ish language, capital letters, a generous amount of DO NOTs and MUSTs Humans are hard to change. But AI agents are not. And that is the only lever we need for better organizationsImage hidden -
-
@bprintco just make a cli for your crm x.com/onusoz/status/… -
-
just added session persistence to our kubernetes managed devboxes using zmx by Eric Bower (neurosnap/zmx on github). like tmux but with native scrollback! I don't want to give agents access to my personal computer, so I host them on hetzner. one click spawn, and start working -
@nicopreme I do something equivalent on codex with just a skill Ralphing works 90% of the time with reviews, and if it gives a stupid review, you just revert -
-
-
-
Here is the project, attaching to multiple sessions is pretty seamless github.com/neurosnap/zmx?…Image hidden -
TIL: zmx session persistence like tmux or gnu screen, but you can scroll up natively! uses @mitchellh's libghostty-vt to attach/restore previous sessions link belowImage hidden -
-
@mazeincoding it’s not the model it’s cursor rate limiting you -
The fundamental problem with GitHub is trust: humans are to be trusted. If you don't trust a human, why did you hire them in the first place? Anyone who reviews and approves PRs bears responsibility. Rulesets exist and can enforce e.g. CODEOWNER reviews or only let certain people make changes to a certain folder But the initial repo setup on GitHub is allow-by-default. Anyone can change anything until they are restricted from it This model breaks fundamentally with agents, who are effectively sleeper cells that will try to delete your repo the moment they encounter a sufficiently powerful adversarial attack For example, I can create a bot account on github and connect @openclaw to it. I need to give it write permission, because I want it to be able to create PRs. However, I don't want it to be able to approve PRs, because a coworker could just nag at the bot until it approves a PR that requires human attention To fix this, you have to bend backwards, like create a @ human team with all human coworkers, make them codeowner on /, and enforce codeowner reviews. This is stupid and there has to be another way Even worse, this bot could be given internet access and end up on a @elder_plinius prompt hack while googling, and start messing up whatever it can in your organization It is clear that github needs to create a second-class entity for agents which are default low-trust mode, starting from a point of least privilege instead of the other way around -
STOP using Claude Code and Sl(opus) to code if ❌ you are not a developer, ❌ or you are an inexperienced dev, ❌ or you are an experienced dev but working on a codebase you don't understand If you *are* any of these, then STOP using models that are NOT state of the art. (See below for what you *should* use) When you don't know what you are doing, then at least the model should know what you are doing. The less knowledgeable and opinionated you are, the more knowledgeable and smart the AI has to be In other words, the AI has to compensate for your deficiencies. Always pay for the best AI you can. It will save you time AND money (thanks to lower token usage and better one-shotting) You pay MORE to pay LESS. It is paradoxical, I know, but it is also proven, e.g. when Sonnet ends up using more tokens than Slopus and ends up costing higher, because it has to try many times more 👨🏻⚕️ For January 2026, your family engineer recommends GPT 5.2 Codex with Extra High Reasoning for general usage and vibe coding. IMPORTANT: Not medium. Not high. EXTRA high reasoning When you use it, you will notice that it is SLOW. Can you guess why? Because it is THINKING more. So it doesn't make the mistakes Slopus makes. This way, you can spend the time handholding a worse model to instead step back and multi-task on some other task and create 3-5x more work The state of the art will most likely change in one month. Don't get married to a a model... There is no loyalty in AI... The moment a better model comes, I will ditch the old one and use that one. I am on the part of this sector that is trying to reduce switching costs to zero I can't wait until I get GPT 5.2 xhigh level of quality with open models, and for 100x cheaper and faster! Until then, make sure to try every option and choose the one that is most reliable for you Follow me to get notified when a new SOTA drops for agentic engineering -
Codex agrees. Sycophant pehImage hidden -
-
It is clear at this point is that github's trust and data models will have to change fundamentally to accommodate agentic workflows, or risk being replaced by other SCM One *cannot* do these things easily with github now: - granular control: this agent running in this sandbox can only push to this specific branch. If an agent runs amok, it could delete everybody's branches and close PRs. github allows for recovery of these, but still inconvenient even if it happens once - create a bot (exists already), but remove reviewing rights from it so that an employee cannot bypass reviews by tricking the bot to approve - in general make a distinction between HUMAN and AGENT so that you can create rulesets to govern the relationships in between cc @jaredpalmer -
-
Automated AI reviews on github by creating an ai-review skill and a script to paste trigger prompts and wait for their response. It is instructed to loop and not stop until all AI review feedback is resolved. This AI review workflow developed gradually based on the current capabilities, and I've realized recently that it became quite mechanical. So decided to automate it in full ralph spirit (it's ok because it's addressing feedbacks and fixing minor bugs) In the current state, we paste the contents of REVIEW_PROMPT.md into a comment, which automatically tags claude (opus 4.5) and codex (whatever model openai is serving) It then waits until both have responded. In the ai-review skill, it is instructed to take the feedback from SLopus with a grain of salt and ignore feedback that doesn't make sense It works! See in the images below. If the review is stupid, you will of course see it on the PR and what the model has done, and can revert itImage hiddenImage hidden -
GitHub has to change
It is clear at this point is that GitHub’s trust and data models will have to change fundamentally to accommodate agentic workflows, or risk being replaced by other SCM
One cannot do these things easily with GitHub now:
- granular control: this agent running in this sandbox can only push to this specific branch. If an agent runs amok, it could delete everybody’s branches and close PRs. GitHub allows for recovery of these, but still inconvenient even if it happens once
- create a bot (exists already), but remove reviewing rights from it so that an employee cannot bypass reviews by tricking the bot to approve
- in general make a distinction between HUMAN and AGENT so that you can create rulesets to govern the relationships in between
The fundamental problem with GitHub is trust: humans are to be trusted. If you don’t trust a human, why did you hire them in the first place?
Anyone who reviews and approves PRs bears responsibility. Rulesets exist and can enforce e.g. CODEOWNER reviews or only let certain people make changes to a certain folder
But the initial repo setup on GitHub is allow-by-default. Anyone can change anything until they are restricted from it
This model breaks fundamentally with agents, who are effectively sleeper cells that will try to delete your repo the moment they encounter a sufficiently powerful adversarial attack
For example, I can create a bot account on GitHub and connect clawdbot to it. I need to give it write permission, because I want it to be able to create PRs. However, I don’t want it to be able to approve PRs, because a coworker could just nag at the bot until it approves a PR that requires human attention
To fix this, you have to bend backwards, like create a @human team with all human coworkers, make them codeowner on /, and enforce codeowner reviews. This is stupid and there has to be another way
Even worse, this bot could be given internet access and end up on a @elder_plinius prompt hack while googling, and start messing up whatever it can in your organization
It is clear that GitHub needs to create a second-class entity for agents which are default low-trust mode, starting from a point of least privilege instead of the other way around
-
Now it’s Claude Code’s turn to implement queueing -
-
Codex users rejoice Also, pi is officially not shitty: shittycodingagent. ai -> buildwithpi. ai since a few days -
with ai, writing correct tests is now the bottleneck in projects like this web-platform-tests are already there now let’s see if someone will beat @ladybirdbrowser to it -
As someone who is frontrunning mainstream by roughly 6 months, I can tell you that you will be raving about pi and @openclaw 6 months instead of claude code. Go check them out at clawd.bot and shittycodingagent.ai -
Kullanmayan agent’ı, alamaz maaşı -
-
I propose a new way to distribute agent skills: like --help, a new CLI flag convention --skill should let agents list and install skills bundled with CLI tools Skills are just folders so calling --skill export my-skill on a tool could just output a tarball of the skill. I then set up the skillflag npm package so that you can pipe that into: ... | npx skillflag install --agent codex which installs the skill into codex, or any CLI tool you prefer. Supports listing skills bundled with the CLI, so your agents know exactly what to installImage hidden -
You don't need a skill registry (for your CLI tools)
tl;dr I propose a CLI flag convention
--skilllike--helpfor distributing skills and try to convice you why it is better than using 3rd party registries. See osolmaz/skillflag on GitHub.
MCP is dead, long live Agent Skills. At least for local coding agents.
Mario Zechner has been making the point that CLI tools perform better than MCP servers since a few months already, and in mid December Anthropic christened skills by launching agentskills.io.
They had introduced the mechanism to Claude Code earlier, and this time they didn’t make the mistake of waiting for OpenAI to make a provider agnostic version of it.
Agent skills are basically glorified manpages or
--helpfor AI agents. You ship a markdown instruction manual inSKILL.mdand the name of the folder that contains it becomes an identifier for that skill:my-skill/ ├── SKILL.md # Required: instructions + metadata ├── scripts/ # Optional: executable code ├── references/ # Optional: documentation └── assets/ # Optional: templates, resourcesPossibly the biggest use case for skills is teaching your agent how to use a certain CLI you have created, maybe a wrapper around some API, which unlike
gh,gcloudetc. will never be significant enough to be represented in AI training datasets. For example, you could have created an unofficial CLI for Twitter/X, and there might still be some months/years until it is scraped enough for models to know how to call it. Not to worry, agent skills to the rescue!Anthropic, while laying out the standard, intentionally kept it as simple as possible. The only assertions are the filename
SKILL.md, the YAML metadata, and the fact that all relevant files are grouped in a folder. It does not impose anything on how they should be packaged or distributed.This is a good thing! Nobody knows the right way to distribute skills at launch. So various stakeholders can come up with their own ways, and the best one can win in the long term. The more simple a standard, the more likely it is to survive.
Here, I made some generalizing claims. Not all skills have to be about using CLI tool, nor most CLI tools bundle a skill yet. But here is my gut feeling: the most useful skills, the ones worth distributing, are generally about using a CLI tool. Or better, even if they don’t ship a CLI yet, they should.
So here is the hill I’m ready to die on: All major CLI tools (including the UNIX ones we are already familiar with), should bundle skills in one way or another. Not because the models of today need to learn how to call
ls,greporcurl—they already know them inside out. No, the reason is something else: establish a convention, and acknowledge the existence of another type of intelligence that is using our machines now.There is a reason why we cannot afford to let the models just run
--helporman <tool>, and that is time, and money. The average--helpor manpage is devoid of examples, and is written in a way thay requires multiple passes to connect the pieces on how to use that thing.Each token wasted trying to guess the right way to call a tool or API costs real money, and unlike human developer effort, we can measure exactly how inefficent some documentation is by looking at how many steps of trial and error a model had to make.
Not that human attention is less valuable than AI attention, it is more so. But there has never been a way to quantify a task’s difficulty as perfectly as we can with AI, so we programmers have historically caved in to obscurantism and a weird pride in making things more difficult than they should be, like some feudal artisan. This is perhaps best captured in the spirit of Stack Overflow and its infamous treatment of noob questions. Sacred knowledge shall be bestowed only once you have suffered long enough.
Ahh, but we don’t treat AI that way, do we? We handhold it like a baby, we nourish it with examples, we do our best to explain things all so that it “one shots” the right tool call. Because if it doesn’t, we pay more in LLM costs or time. It’s ironic that we are documenting for AI like we are teaching primary schoolers, but the average human manpage looks like a robot novella.
To reiterate, the reason for this is two different types of intelligences, and expectations from them:
- An LLM is still not considered “general intelligence”, so they work better by mimicking or extending already working examples.
- A LLM-based AI agent deployed in some context is expected to “work” out of the box without any hiccups.
On the other hand,
- a human is considered general intelligence, can learn from more sparse signals and better adapt to out of distribution data. When given an extremely terse
--helpor manpage, a human is likelier to perform better by trial and error and reasoning, if one could ever draw such a comparison. - A human, much less a commodity compared to an LLM, has less pressure to do the right thing every time all the time, and can afford to do mistakes and spend more time learning.
And this is the main point of my argument. These different types of intelligences read different types of documentation, to perform maximally in their own ways. Whereas I haven’t witnessed a new addition to POSIX flag conventions in my 15 years of programming, we are witnessing unprecedented times. So maybe even UNIX can yet change.
To this end, I introduce
skillflag, a new CLI flag convention:# list skills the tool can export <tool> --skill list # show a single skill’s metadata <tool> --skill show <id> # install into Codex user skills <tool> --skill export <id> | npx skillflag install --agent codex # install into Claude project skills <tool> --skill export <id> | npx skillflag install --agent claude --scope repoClick here for the repo, osolmaz/skillflag on GitHub
For example, suppose that you have installed a CLI tool to control Philips Hue lights at home,
hue-cli.To list the skills that the tool can export, you can run:
$ hue-cli --skill list philips-hue Control Philips Hue lights in the terminalYou can then install it to your preferred coding agent, such as Claude Code:
$ hue-cli --skill export philips-hue | npx skillflag install --agent claude Installed skill philips-hue to .claude/skills/philips-hueYou can optionally install the skill to
~/.claude, to make it global across repos:$ hue-cli --skill export philips-hue | npx skillflag install --agent claude --scope user Installed skill philips-hue to ~/.claude/skills/philips-hueOnce this convention becomes commonplace, agents will by default do all these before they even run the tool. So when you ask it to “install hue-cli”, it will know to run
--skill listthe same way a human would run--helpafter downloading a program, and install the necessary skills themselves without being asked to. -
Anthropic earlier last year announced this pricing scheme $20 -> 1x usage $100 -> 5x usage $200 -> 1̶0̶x̶ 20x usage As you can see, it's not growing linearly. This is classic Jensen "the more you buy, the more you save" But here is the thing. You are not selling hardware like Jensen. You are selling a software service *through an API*. It's the worst possible pricing for the category of product. Long term, people will game the hell out of your offering Meanwhile OpenAI decided not to do that. There is no quirky incentive for buying bigger plans. $200 chatgpt = 10 x $20 chatgpt, roughly And here is where it gets funny. Despite not having such an incentive, you can get A LOT MORE usage from the $200 OpenAI plan, than the $200 Anthropic plan. Presumably because OpenAI has better unit economics (sama mentioned they are turning a profit on inference, if you are to believe) Thanks to sounder pricing, OpenAI can do exactly what Anthropic cannot: offer GPT in 3rd party harnesses and win the ecosystem race Anthropic has cornered itself with this pricing. They need to change it, but not sure if they can afford to do so in such short notice All this is extremely bullish on open source 3rd party harnesses, @opencode, @badlogicgames's pi and such. It is clear developers want options. "Just give me the API" I personally am extremely excited for 2026. We'll get open models on par with today's proprietary models, and can finally run truly sovereign personal AI agents, for much cheaper than what we are already paying!Image hidden -
Anthropic's pricing is stupid
Anthropic earlier last year announced this pricing scheme
- $20 -> 1x usage
- $100 -> 5x usage
- $200 -> 1̶0̶x̶ 20x usage
As you can see, it’s not growing linearly. This is classic Jensen “the more you buy, the more you save”
But here is the thing. You are not selling hardware like Jensen. You are selling a software service through an API. It’s the worst possible pricing for the category of product. Long term, people will game the hell out of your offering
Meanwhile OpenAI decided not to do that. There is no quirky incentive for buying bigger plans. $200 chatgpt = 10 x $20 chatgpt, roughly
And here is where it gets funny. Despite not having such an incentive, you can get A LOT MORE usage from the $200 OpenAI plan, than the $200 Anthropic plan. Presumably because OpenAI has better unit economics (sama mentioned they are turning a profit on inference, if you are to believe)
Thanks to sounder pricing, OpenAI can do exactly what Anthropic cannot: offer GPT in 3rd party harnesses and win the ecosystem race
Anthropic has cornered itself with this pricing. They need to change it, but not sure if they can afford to do so in such short notice
All this is extremely bullish on open source 3rd party harnesses, OpenCode, Mario Zechner’s pi and such. It is clear developers want options. “Just give me the API”
I personally am extremely excited for 2026. We’ll get open models on par with today’s proprietary models, and can finally run truly sovereign personal AI agents, for much cheaper than what we are already paying!
-
The models, they just wanna work. They want to build your product, fix your bugs, serve your users. You feed them the right context, give them good tools. You don’t assume what they cannot do without trying, and you don’t prematurely constrain them into deterministic workflows. -
-
-
-
This, and insisting on CLAUE.md are really lame @AnthropicAI -
-
-
-
-
-
-
-
-
.@openclaw workspace and memory files can be version-controlled! In our pod, inotify triggers a watcher script every time there is a change to workspace folder, to sync these files to our monorepo. It then goes through the same steps: - Create zeno-workspace branch if doesn't exist, otherwise, skip - Sync changes to the branch, then commit - Create PR on github if doesn't exist - PRs can then be merged every once in a while, after accumulating enough changes. Merge triggers re-deploy, and clawd restarts with the same state Simple foolproof automatic persistence for remote CI/CD handled clawd (except for when you are running multiple clawds at the same time, but we are not there yet) cc @steipeteImage hiddenImage hidden -
-
-
pi now supports your openai plus/pro subscription -
-
Having a "tools" repo as a developer
I am a fan of monorepos. Creating subdirectories in a single repo is the most convenient way to work on a project. Low complexity, and your agents get access to everything that they need.
Since May 2025, I have been increasingly using AI models to write code, and have noticed a new tendency:
- I don’t shrug from vendoring open source libraries and modifying them.
- I create personal CLIs and tools for myself, when something is not available as a package.
With agents, it’s really trivial to say “create a CLI that does X”. For example, I wanted to make my terminal screenshots have equal padding and erase cropped lines. I created a CLI for it, without writing a single line of code, by asking Codex to read its output and iterate on the code until it gives the result I wanted.
Most of these tools don’t deserve their own repos, or deserve being published as a package at the beginning. They might evolve into something more substantial over time. But at the beginning, they are not worth creating a separate repo for.
To prevent overhead, I developed a new convention. I just put them in the same repo, called tools. Every tool starts in that repo by default. If they prove themselves overly useful and I decide to publish them as a package, I move them to a separate repo.
You can keep
toolspublic or private, or have both a public and private version. Mine is public, feel free to steal ones that you find useful. -
@rauchg indeed -
75k lines of Rust later, here is what I’ve built during the first Christmas with agents, using OpenAI Codex 🎄🤖 - A full mobile rewrite and port of my Python Instagram video production pipeline (single video production time: 1hr -> 5min) (ig: nerdonbars) - Bespoke animation engine using primitives (think Adobe Flash, Manim) - Proprietary new canvas UI library in Rust, because I don’t want to lock myself into Swift - Thanks to that, it’s cross platform, runs both on desktop and iOS. It will be a breeze porting this to Android when the time comes - A Rust port of OpenCV CSRT algorithm, for tracking points/objects - In-engine font rendering using rustybuzz, so fonts render the same everywhere - Many other such things Why would I choose to do it that way? Because I have developed it primarily on desktop where I have much faster iteration speed. Aint nobody got time for iOS compilation and simulator. Once I finished the hard part on desktop, porting to iOS was much easier, and I didn’t lock myself in to Apple Some of these would have been unimaginable without agents, like creating a UI library from scratch in Rust. But when you have infinite workforce, you can ask for crazy things like “create a textbox component from scratch” What I’ve built is very similar in nature to CapCut, except that I am a single person and I’ve built it over 1 week What have you built this Christmas with agents? cc @thsottiaux -
-
Christmas of Agents
I believe a “Christmas of Agents” (+ New Year of Agents) is superior to “Advent of Code”.
Reason is simple. Most of us are employed. Advent of Code coincides with work time, so you can’t really immerse yourself in a side project.1
However, Christmas (or any other long holiday without primary duties) is a better time to immerse yourself in a side project.
2025 was the eve of agentic coding. This was the first holiday where I had full credential to go nuts on a side project using agents. It was epic:
Tweet embed disabled to avoid requests to X.75k lines of Rust later, here is what I’ve built during the first Christmas with agents, using OpenAI Codex
- A full mobile rewrite and port of my Python Instagram video production pipeline (single video production time: 1hr -> 5min)
- Bespoke animation engine using primitives (think Adobe Flash, Manim)
- Proprietary new canvas UI library in Rust, because I don’t want to lock myself into Swift
- Thanks to that, it’s cross platform, runs both on desktop and iOS. It will be a breeze porting this to Android when the time comes
- A Rust port of OpenCV CSRT algorithm, for tracking points/objects
- In-engine font rendering using rustybuzz, so fonts render the same everywhere
- Many other such things
Why would I choose to do it that way? Because I have developed it primarily on desktop where I have much faster iteration speed. Aint nobody got time for iOS compilation and simulator. Once I finished the hard part on desktop, porting to iOS was much easier, and I didn’t lock myself in to Apple
Some of these would have been unimaginable without agents, like creating a UI library from scratch in Rust. But when you have infinite workforce, you can ask for crazy things like “create a textbox component from scratch”
What I’ve built is very similar in nature to CapCut, except that I am a single person and I’ve built it over 1 week
What have you built this Christmas with agents?
-
You could maybe work in the evening after work, but unless you are slacking at work full time, it won’t be the same thing as full immersion. ↩
-
-
Migrating @TextCortex to SimpleDoc. It's really easy with the CLI wizard! npx @simpledoc/simpledoc migrate We have a LOT of docs spanning back to 2022, pre coding agent era. Now we will have CI/CD in place so that coding agents can't litter the repo with random Markdown filesImage hidden