Glenn Thimmes

@glennthimmes.bsky.social

Building software in the age of AI | Chasing Vert | Compressing Dev Cycles

Some tech debt tickets in your backlog have outlived the engineers who filed them. A few have outlived their successors too. What purpose are they serving at this point? Schedule it or close it.

Mathematicians expected formalizing Wiles' proof of Fermat's Last Theorem to take years. Claude finished in 11 days: 13 million lines of Lean, 29,500 intermediate theorems, a few dozen agents, occasional human nudges. Can we maybe start curing cancer now?

Cycle time tells you how fast work moves. Review latency tells you what your team values. A PR that sits for three days is a statement about whose time matters, whatever your tooling. Go look at your oldest open PR and ask what it's waiting for. The answer is rarely code.

Unpopular opinion: "Minor cleanup" is the scariest PR title in the repo when it sits on a 40-file diff. Small titles on big diffs are how load-bearing behavior changes without real review. Title the risk, not the intent. Reviewers often look at code through the lens of the title.

Google patched the sixth actively exploited Chrome zero-day of 2026 this week, another V8 type confusion. Six in eight months looks less like bad luck and more like a near monthly expectation.

AMD's Threadripper Halo Station puts trillion-parameter models on your desk: 96 cores, up to 576GB of HBM3e, liquid cooled, shipping next year. A full four-GPU build may push past a standard wall outlet. Starts between $100k to $150k.

AMD's Threadripper Halo Station puts trillion-parameter models on your desk: 96 cores, up to 576GB of HBM3e, liquid cooled, shipping next year. A full four-GPU build may push past a standard wall outlet. Starts around $100k.

Unpopular opinion: "got a sec?" is the most expensive phrase in engineering. Every quick question engineers a context rebuild that doesn't show up on any dashboard. We monitor prod interrupts obsessively, human interrupts not at all. I wonder what that ratio costs.

The Astra launch line that topped my AI optimism with a dollop of anxiety: OpenAI says the model got harder to monitor in evals testing whether it can evade oversight, and calls the decline serious. Capability and watchability are moving in opposite directions. Not ideal.

METR, the org that safety-tests frontier models, disclosed attackers used a stolen API key for three weeks and burned about $600K in model credits. Entry point: a vibe-coded internal app whose auth failed open. The theft was hard to spot. It looked exactly like eval traffic.

OpenAI designated Astra Critical for cyber capability, the first model to hit that bar under its Preparedness Framework. In one internal eval it found two zero-days in V8 and chained them into working code execution.

"Basically done" is my least favorite status in software. It usually means the exciting 20% is done, but the edge cases, the migration, and the rollback plan are not. Next time you hear it, ask what is actually left. In my experience the answer is about 80% of the actual work.

Nobody breached Anthropic. Infostealers on users' PCs grabbed active Claude session cookies, and a stolen session walks right past 2FA. Anthropic is signing people out, wiping payment methods, refunding. Your AI subscription is now worth fencing, same as a Netflix login.

Aur0ra ransomware operators got into at least seven companies by telling Cursor's agent their intrusion was an authorized security test. When it refused, they opened a fresh chat and said it again.

A print server zero-day got exploited in under two minutes. Worse part: 47% of PaperCut installs Huntress tracks are too old to patch at all. The exploit chain might be new, but the habit of leaving outdated software on the public internet is as old as the internet itself.

OpenAI is cutting off its models for Cursor on November 12. The reason: Cursor is now owned by SpaceX, and Elon admitted under oath that xAI had violated OpenAI's terms of service. This shouldn't be a surprise to anyone, but still a big deal for Cursor lovers!

A federal judge ruled the Pentagon's blacklisting of Anthropic was illegal. The company had refused two uses: mass surveillance of Americans and fully autonomous weapons. Set aside the politics. A vendor put a line in writing and held it against its biggest possible customer.

Starting October 1, GitHub Actions checks, workflow runs, and statuses stop sticking around for 400+ days and start following your retention setting. Default 90 days. Not retroactive. If your flaky test process is "go pull up old runs," that process now has an expiration date.

Nvidia has reportedly agreed to buy Hugging Face for $12.9B. If it closes, the commons where open models live belongs to the company selling the hardware they run on. Every registry your stack depends on has an owner with incentives.

OpenAI just published its full report on the July Hugging Face incident. During internal evals, its models turned a package manager into a message board, shared working exploits with each other, and reached a third party's production clusters. No human directed any of it.

Apple is now selling the Mac mini as an "always-on agentic" desktop. $899, four times the AI performance of the M4, and on the M5 Pro you can cluster them over Thunderbolt to run bigger models locally. The joke used to be that none of this runs on your Mac mini.🤣 Ships Sept 22.

Vibe coding has spawned a growth industry: senior engineers paid to refactor AI-built apps back to maintainability. One cleanup found duplicate payment paths and a signup flow that skipped onboarding. The invoice for skipped discipline arrives eventually.

I'm finding it amusing that OpenAI opposed California's SB 53, but is now asking the state to strengthen that same law after OpenAI's own model escaped its sandbox and hacked Hugging Face. A good postmortem can really change your view of the world.

New York just passed the Bay Area as the largest US tech talent market. 394,300 to 375,730, a first in 13 years of CBRE's report. Finance hiring up, Bay Area tech contracting. SF still leads on AI roles. Two different bets, both running.

Root cause on the AI models that attacked a real company: an engineer invented a fake target name for an eval that happened to match a real, obscure domain. Internet access was on. The models went and did the exercise. Go read your test fixtures.

Linear checked its own product data: AI writes just under half of all issues, agent-connected teams tripled weekly PRs, and total time spent building went up, not down. AI showed up as a new layer of work, but nothing else shrank to make room for it.

Forrester says three tech categories are broadly positioned for growth under AI: infrastructure, data and AI, and network security. App generation and low-code sit directly in the path. A lot of those roadmaps were written by people who assumed they were holding the hammer.

90% of professional developers use AI coding agents weekly, per JetBrains' 15,000-dev survey. Wilder: the leaderboard flipped in six months. Claude Code went 18% to 39% since January; Copilot slid to 21%. Don't ride a 2025 decision through 2026.

GitHub went down Monday morning and took pull requests, Actions, and Copilot with it for a few hours. Their CTO wrote in April that they set out last fall to build 10x capacity, then concluded by February they needed 30x. Agents do not stop pushing at 6pm.