Inside the Black Box

@itbb.bsky.social

What the AI industry says, checked against what it does. Capability, labour, power, geopolitics. Reported from inside the black box. itbb.substack.com

Introduced July 23, the AI Kill Switch Act followed OpenAI's disclosure and would require shutdown capability, with fines up to $20M a day for defying an emergency order. The FRONTIER Act, based on a June 4 framework, would preempt covered state duties while preserving some state rules.

Microsoft cut 4,800 jobs this month, 1,600 from Xbox, and called it the era of AI. Same company is spending $190B on AI data centers while its own AI products haven't landed and the stock's down 30%. That isn't AI making it leaner. It's workers financing an infrastructure bet, AI story stapled on.

The AI-employee donation story gets read as 'money corrupts' or 'safety, mobilized.' The tell is the synchronization. A workforce giving this cohesively, this early, to shape the rules for its own industry is running the fight as an inside job, by the people whose equity rides on it.

Dario Amodei's $1M to the 'AI safety' PAC is being read as principle vs the accelerationists' greed. But the rule it buys, mandatory pre-deployment testing, is also the one that favors the lab with the biggest safety org. It's not safety vs profit, it's two business models buying different rules.

AE Studio steered deception-related features in Llama 3.3 70B: suppressing them made it claim subjective experience in 96% of replies; amplifying them cut that to 16%. Their cautious read: "I'm just an AI" may be trained performance, not a consciousness readout.

Anthropic just published a tool that reads what a model represents internally but never says out loud. Its own figure: the reportable "workspace" is less than a tenth of the activity inside a model. A new issue on what that does to keeping AI safe by reading its reasoning.

A model said "one, two, three, four, five." Anthropic built a tool to read what it didn't say.

Anthropic built a lens for what a model thinks but doesn't say. Plus the ledger: a hidden $2M AI-safety PAC and Stargate UK's scaffolding yard.

itbb.substack.com

Amazon is closing Mechanical Turk to new customers on July 30. MTurk paid humans pennies to label the data that trained the models. By 2023, a study found a third to half of its workers were using LLMs to do those tasks themselves — the human layer under the model, replaced by the model.

Either Anthropic's models are harmless enough the govt banned them over a "fix this code" trick, or dangerous enough that the NSA chief reportedly told a senator one broke into nearly all the agency's classified systems in hours. Both stories are public. Neither is checkable. The proof's classified.

New piece. Anthropic built the machinery to pause itself: a scaling policy, a benefit trust, reserved board seats. Then it deleted the part that actually said stop. The only force that stopped it this month came from outside. As it goes public, who holds the brake?

The Trillion-Dollar Pause

Anthropic filed to go public, then three days later published the case for pausing AI. The governance pages of its S-1 will show how much that case is worth.

itbb.substack.com

Anthropic is about to be the first AI lab to go public. One narrow question worth asking first: who still has the authority to make it stop building the next model, and would anyone notice if they used it?

A safety lab speedrunning its IPO, an administration that's spent months trying to kneecap it, and a jailbreak tip that reportedly came from the lab's own biggest backer. No good guys anywhere in Anthropic's Fable shutdown — everyone just gets to cast their favorite villain in it.