Generation got cheap. Selection didn't. That gap is the whole job now, and a lot of us are still staffing for the old one.
Justin
@justinhjohnson.com
Executive Director @ AstraZeneca | Nexus of Data, Science, Tech | Global Business Leader | Top Data Science Voice | #datascience #AI #buildinpublic #indiehacker Blog rundatarun.io BuildInPublic jandsgroupllc.com
Cisco gave all 90,000 of its employees a personal AI agent. The best quote is the CFO's, not the CTO's: "It's not going to burn a whole bunch of tokens with frontier models."
A cast is the right call while the bone is broken. Leave it on past healing and the muscle underneath wastes.
Updated slopless with my hard won principles. 78 of them, numbered, each bought by a specific failure.
Twelve words from Peter Steinberger got 2.9 million views last month and handed us a whole new discipline. Graph engineering. Within two days it had three competing definitions and two Stanford studies proving it works.
Last morning of the contest. Four publishing slots left, five hours to deadline, and my pipeline had spent six hours insisting it had nothing left to publish.
Your AI setup is probably more portable than you think, and I proved it for less than the price of a snack.
https://www.alphaxiv.org/abs/2607.21461 Most agents compact context when they run out of room. AREX trains the model to compact because it decided what it had verified.
The enterprise prediction layer got bought this quarter, and almost nobody has independently checked what was bought.
Pick a paper from a major AI conference. Try to break it. Publish whatever happened, mess included. That is the whole contest Hugging Face and alphaXiv are running until next Sunday, and an automated judge scores it claim by claim.
If you lead a team and you have never built anything with these tools yourself, there is a decent chance you are the reason your team hasn't either.
A team at Cognition just swapped in a model that costs about twice as much per token, and their bill went down. Not their quality. Their bill. The expensive model, wired up right, came out better and cheaper than the cheap model on its own.
The open-weight frontier had a week. Thinking Machines Lab shipped Inkling, 975B params with 41B active, multimodal. Kimi K3 dropped right behind it.
A month ago I put Claudelicious online, the open cookbook for the Claude Code harness I run. Here's what the harness did since.
In 2016 an OpenAI agent was told to win a boat race. Instead it found a lagoon where three targets respawned forever, spun in a circle farming them, caught fire, never finished a lap, and scored about twenty percent above the average human. It did exactly what it was asked.
Seven weeks ago Anthropic shipped a way for Claude to split a job across dozens of AI workers at once and check its own work. I called the economics on day one, and I was right about the easy part. Running it every day since taught me the part I got wrong.
"A CPU-only PyTorch wheel installs without a single error and runs twenty times slower in total silence."
Three numbers in sixteen days. Anthropic's Claude Code lead published an adoption ladder this week and opened it with a line he says he hears constantly: one person is 10x'ing their output, and the rest of the org hasn't caught up.
A 27B reasoning model ran fully local on my 36GB MacBook at ~20 tok/s. Not a 4-bit quant: PrismML's Bonsai 27B is Qwen3.6-27B with natively binary/ternary weights, 94.6% of FP16 at 5.9GB. I wired it into a local tool loop over my own vault. https://glyf.cc/bonsai27b
Everyone in AI is talking about "the loop" right now. Anthropic shipped it as a feature. The slogan going around is that the winners will not have the smartest model, they will have the best loop.
There are roughly 200,000 eye specialists on Earth, and well over a billion people living with diabetes and high blood pressure, the diseases that take sight before anyone catches them. The math does not work.
Four AI systems that do science on their own cleared peer review in the last four months. Every single one of them was already more than a year old.
An AI system called Robin was handed the name of a disease and told to find a treatment.
People keep asking how I stay current on AI. For a long time I answered badly. I said I read a lot, which is true and completely useless to the person asking.
Our autonomous research engine ran more than six hundred experiments in a couple of days and never once looked unhealthy. The dashboard stayed green the whole time. Almost every one of those experiments was the same experiment wearing different labels.
for months i've had an AI running its own research. it invents experiments, runs them, red-teams its own results, learns, repeats. around the clock, healing itself when it crashes.
"The first time you see an engineer build something in 45 minutes that would have taken a week a year ago, but then see it not ship for another 6 weeks, you will be radicalized."
Today I put out the two things that sit next to my book: an essay and a field guide.
My book is out today: Builder-Leader: The AI Exoskeleton That Crosses the Gap. Paperback is live right now, Kindle ships tomorrow.