Hillary Sanders

@meatlearner.bsky.social

Machine-learner, meat-learner, research scientist, AI Safety thinker. Model trainer, skeptical adorer of statistics. Co-author of: Malware Data Science

In light of the OpenAI hacking incident, I’d rather criminal liability rest on the model [family] “itself.” That is: once charged / convicted, it becomes illegal for any person or agent to use/deploy the model & descendants for [AI sentence] years. Probation may require proof of improved safety.

tv.apple.com has this excellent feature where if you go back a few seconds in a show, it temporarily turns on subtitles, as understanding what was said is a common reason for going back. This is surprising, because their overall web user interface design is atrocious.

Today’s frontier models train in an expensive style: dense forward passes, huge matrix multiplies, and broad weight updates. The human brain (~5 MWh over 28 years) is an existence proof that learning can be vastly more energy efficient - about 10,000x - than modern AI training runs.

The 2026 Int'l AI Safety Report says AI progress could accelerate “if AI systems begin to speed up AI research itself.” But that "if" is no longer an if. The question is how strong the feedback loop will be, and how much it will be slowed by bottlenecks & diminishing returns.

Bild

The hardest AI risks to govern are high-severity, low-evidence, fast-moving risks. A misaligned system that can hack, replicate, and seek compute requires swift action. Arguably: large GPU clusters should be internationally tracked.

Hot take: fewer frontier AI labs makes safety regulation easier. For once, less competition is better. Meanwhile... when Anthropic declined DoD uses involving mass surveillance and fully autonomous weapons, the current administration designated it a supply-chain risk. Awful.

Bild

Compersion, love, and the urge to help teach and enable us to stand on our own are traits that are innate to mothers' brains, and should also be strongly present in future super-intelligent AI systems. If we're going to build AGI, let's model it after mothers.

In AI safety, we have inner misalignment (actions don't minimize the loss function) and outer misalignment (loss function is misspecified). But I do think that inner misalignment (~learned features) tend to act as a protective mechanism to avoid outer misalignment implications. I, er, really hope.