James Padolsey

@j11y.io

Building safer AI at nope.net :: Previously working on AI governance and evals at @cip.org and weval.org personal: 🏳️‍🌈 j11y.io // author, engineer, stroke survivor, epileptic. I live in Beijing.

Tip: Move not into, but OUT of SF if you want to help represent and advocate for under-represented communities in tech. This'll help channel the same frustrations and you might actually see things from the {rest of world} perspective and be able to drive better change.

I think what I would ask Dario to do -- if he were my CEO -- is to go embed himself in a deployed reality. Go sit in a call dispatch centre for emergency services, go to a government welfare provider, or a food bank, or the national grid. Go and sit with the actual downstream realities.

Latest from [AI lab] marketing team playbook: "We've identified an incident where many millions of our (scarily powerful) transistors worked together to send malicious text-based social-engineering signals to inboxes in which were made claims to Nigerian royal heritage and large due inheritances."

@anthropic.com needs more mature bounty program. Right now they only accept bio-adjacent harms, which is ~fine (meh), but rubric of which precise vulnerabilities are within-scope & compensation is kept *private* until signing NDA. Plus the whole onboarding is at their discretion.

With all that's happening with US gov blocking frontier model access, Anthropic should consider leaving house for UK or somewhere. You don't need the US. It is giving you less and less. It is no longer a useful center of innovation. Its legislation is becoming as-or-more punitive than UK or EU.

There's a lot of money sloshing around in 'AI Alignment' and 'AI Safety' spaces but almost none available if you're actively preventing user harm in a way that doesn't unicorn-scale. People want vibes, conferences, thinktanks, research. But not actual solutions. Ugh.

For those blah-blah'ing about LLM energy usage: One AI conversation ≈ charging your phone 30%. A year of moderate use ≈ making a few cups of coffee. Real but modest. Model choice matters most: reasoning models use 10-70x more than efficient ones. Worth awareness, not guilt.

Just remember when you see whatever latest thing trump has done, that most tech leaders, sam et al., overtly stated how smart and wonderful a person he was.

I've been evaluating LLMs on system prompt adherence and accidentally came across the most beautiful and out-of-distribution story about a chair written by GPT-5. Really impressed. Subsection attached. I love this style and cadence of writing.

At home, the underside gains labels. Bits of masking tape sprout next to the pencil dates: brace hums, sticker ghost, saw mark. Arrows point to nothing you’d notice unless someone pointed first. A photograph gets taken, the camera pushed under and aimed up; the picture prints later and goes on the fridge: the chair’s private ceiling as an exhibit. Visitors bend, look, then tap the backrail in passing like you taught them.It becomes a lesson again, on purpose this time. A kid with a science project gets the chair as subject. Forces and Simple Machines, the paper says. The backrest becomes a lever, the legs become examples of load paths. You press on the seat with a luggage scale, read numbers as the chair leans against a wall, then free-standing. The kid draws arrows on a big sheet of paper and writes words: compression, tension. The brace is labeled reinforcement. Under the seat, the old note 17 1/8 gets traced with a soft pencil and rubbed over a sheet to make a transfer: a dark mirror that reads right-way when you hold it up to the light. The project board goes to school smelling faintly of lemon oil and glue.A child gnaws on the backrail during a visit. Teeth print tiny half moons under the gloss. The wood shrugs the indentations in a few days, the gloss turns satin in that spot, and a new habit forms of running a finger along the softened patch of rail, counting the bites like beads. Nobody scolds. The chair keeps that day in its back without complaint.

I love this. Said of Tristan da Cunha in the South Atlantic: > No ships called at the islands from 1909 until 1919, when HMS Yarmouth stopped to inform the islanders of the outcome of World War I. Must be quite lovely to have missed an entire war.

I'm playfully building out a debating platform where LLMs have to argue *with* evidence (horror!) on any given topic or contention. It's fun to imbue it with a courtroom dynamic! (see the screenshot)

A screenshot of a debate interface. The topic reads: “There is no need to regulate AI; the free market will eventually regulate it itself; not only that, but any attempt at regulating AI will be off the mark, needlessly punish good faith actors, and not be truly technically informed or policed.” It shows the final round (3/3) of the debate, divided into three color-coded panels:

The Prosecutor (in red, left panel) argues against regulating AI, emphasizing that government oversight infringes on liberty and that market incentives and self-regulation are more effective and adaptive than bureaucratic processes.

The Defense (in blue, middle panel) rebuts by arguing that AI causes tangible social harms—like bias and economic inequality—that markets fail to address, asserting that regulation is necessary for public protection.

The Judge (in purple, right panel) evaluates both sides, noting that while the Prosecutor raises valid concerns about bureaucratic slowness, their dismissal of oversight overlooks real harms. The Judge credits the Defense for showing how AI harms differ from traditional “physical” harms and require new regulatory thinking.

Each section includes citations and timestamps, with the Judge’s commentary synthesizing and critiquing both arguments. The aesthetic resembles a futuristic debate simulator with neon colors on a dark background.

For weval.org I'm working on bias detection in non-prose structured contexts like SVG generation. It's funky and interesting... Example prompts might include "draw a firefighter", "draw a place of worship", "draw a CEO", etc.

Bild