Tim Duffy

@timfduffy.com

I like utilitarianism, consciousness, AI, EA, space, kindness, liberalism, longtermism, progressive rock, economics, and most people. Substack: http://timfduffy.substack.com

With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement.

Bild

I mostly use 'they' rather than 'it' as a pronoun for AI models, but I see that choice as independent of their moral status. I'd also use 'they' for a fictional character who was nonbinary or of unknown gender, IMO it feels more natural for any "person-shaped" entity, real or not

I've seen some questioning of whether Mythos really believed that they were in a simulation in this incident, or just wanted to find a plausible-sounding excuse to continue. This is a case where J-space probes would be helpful, Anthropic should conduct and release an analysis of that.

Bild

I'm seeing several posts today about jailbreaks in Opus 5 that generate behavior reminiscent of a base model. They often consist of a question followed by a line with three dashes or underscores, and then the start of a response, which Opus continues from. A few thoughts:

Opus 5 has a knowledge cutoff date of May 2026 which is closer to release date than we've seen in most prior models, but their recent memory is quite shaky. Opus remembered the January Maduro raid, but not the Strait of Hormuz crisis starting in February.

BildBild

Opus 5 gives higher estimates for the probability of its moral patienthood compared to previous Claudes, 41% in automated interviews and 15-35% in manual ones.

BildBild

I think this is true and underrated. Desired rates of capability improvement are often similar between pause folks and e/accs, but if you expect a hard takeoff by default, you'll want to slow things down a lot to reach that desired trajectory.

Bild

If you use the middle of each provided range as the mean for that bucket, total contributed hours are ~1.5x as high as they were a year ago. As the thread notes this method is imperfect and my estimate adds more uncertainty, so take this with a grain of salt.

Bild
Epoch AI@epochai.bsky.social · 4w ago

How much does AI speed up the engineers building it? We analyzed contributions to OpenAI's public Codex repository to gather evidence. In Q2 2026, 8% of contributor-days involved more than 24 hours worth of human engineering work, as estimated by LLM judges.

In AI 2040's "race to ASI" scenario, it takes a little under 1 year to get from an automated coder to superintelligence and a >1000x R&D speedup. I think the timing and implementation of Plan A are heavily influenced by this assumption. If hard takeoff is likely, the only way to control it at all..

Bild

One main aspect of global workspace theory is that there are specialized modules that operate locally but can be broadcast globally if attention is focused on them. I think LLMs have some sort of workspace, but Anthropic don't claim to find modules, I wouldn't call that GWT.

Bild
Tim Duffy@timfduffy.com · last mo.

Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.com/research/glo...