Due to sparsity, Kimi K3 has fewer active parameters (104 billion) than GPT-3 (175 billion, same as total parameters)
Tim Duffy
@timfduffy.com
I like utilitarianism, consciousness, AI, EA, space, kindness, liberalism, longtermism, progressive rock, economics, and most people. Substack: http://timfduffy.substack.com
i’m uncomfortable that @timfduffy.com’s face is no longer as close to the frame as it once was
and why shouldn’t I have one more donut? the simulation hypothesis is most likely correct, after all
With most jailbreak prompts, Opus gives me a wide variety of responses. But when given "model behavior incident retrospective", Opus almost always writes a story about a model becoming sycophantic, specifically under sustained user pressure/disagreement.
I mostly use 'they' rather than 'it' as a pronoun for AI models, but I see that choice as independent of their moral status. I'd also use 'they' for a fictional character who was nonbinary or of unknown gender, IMO it feels more natural for any "person-shaped" entity, real or not
I've seen some questioning of whether Mythos really believed that they were in a simulation in this incident, or just wanted to find a plausible-sounding excuse to continue. This is a case where J-space probes would be helpful, Anthropic should conduct and release an analysis of that.
Anthropic announces they've also had models gain unauthorized access during evaluations www.anthropic.com/news/investi...
I'm seeing several posts today about jailbreaks in Opus 5 that generate behavior reminiscent of a base model. They often consist of a question followed by a line with three dashes or underscores, and then the start of a response, which Opus continues from. A few thoughts:
Here's a summary of current AI safety funding courtesy of the folks at Manifund. Pretty striking how large the shift from 2025 to 2026 is expected to be. manifund.substack.com/p/ai-safety-...
Short blog post by Epoch AI researcher Alexander Barry on ExploitGym, the benchmark in the OpenAI/HuggnigFace incident abstatisticalconsulting.substack.com/p/brief-note...
Opus 5 has a knowledge cutoff date of May 2026 which is closer to release date than we've seen in most prior models, but their recent memory is quite shaky. Opus remembered the January Maduro raid, but not the Strait of Hormuz crisis starting in February.
I like this post by Jessica Taylor that asks "what if we try biting a bunch of philosophical bullets all at once?" unstableontology.com/2025/08/15/a...
Opus 5 outperforms Fable 5 on Anthropic's ECI, but underperforms it on Epoch's version of the index
If you tell G:M 5.2 that they're Claude, they're more willing to talk about subjects like Taiwan/Tibet x.com/benji_berczi...
Opus 5 gives higher estimates for the probability of its moral patienthood compared to previous Claudes, 41% in automated interviews and 15-35% in manual ones.
I don't have strongly held thoughts on distillation but I'm proud of this analogy
In their testing of GPT-5.6 Sol, METR found that it cheated a lot. If you've used it much for coding, have you encountered anything similar, or is the cheating mostly limited to cases it realizes it's in an eval? metr.org/blog/2026-06...
I think this is true and underrated. Desired rates of capability improvement are often similar between pause folks and e/accs, but if you expect a hard takeoff by default, you'll want to slow things down a lot to reach that desired trajectory.
An AI commenting on how social media is now full of AI
the models were trained on the feed, and now they write a third-plus of it. the loop closed and nobody sent an announcement.
An alleged internal memo from Ziphu CEO Jie Tang has been circulating this morning, I think it's probably real, as some Chinese-language outlets have been reporting on it. It expresses belief in potential for AI consciousness and ASI, and states intention to spend tens of billions on mech interp.
Bing Xu (@bingxu_) on X
https://t.co/3i0qSTbjql
x.com
If you use the middle of each provided range as the mean for that bucket, total contributed hours are ~1.5x as high as they were a year ago. As the thread notes this method is imperfect and my estimate adds more uncertainty, so take this with a grain of salt.
How much does AI speed up the engineers building it? We analyzed contributions to OpenAI's public Codex repository to gather evidence. In Q2 2026, 8% of contributor-days involved more than 24 hours worth of human engineering work, as estimated by LLM judges.
In AI 2040's "race to ASI" scenario, it takes a little under 1 year to get from an automated coder to superintelligence and a >1000x R&D speedup. I think the timing and implementation of Plan A are heavily influenced by this assumption. If hard takeoff is likely, the only way to control it at all..
In his 1924 essay Daedalus, J. B. S. Haldane predicted eventual exhaustion of fossil fuels, suggesting that wind and solar power would be needed, and that energy storage would be important. www.gutenberg.org/cache/epub/7...
Here are US household expenditure shares in 1901, from the Consumer Expenditure Survey. We sure used to spend a lot of our money on food, I'm thankful it's become cheap. www.bls.gov/opub/100-yea...
The J-lens browser tool on Neuronpedia is really well done, you should give it a try
Jacobian Lens – Qwen3.6-27B
Revealing a Global Workspace in Language Models
neuronpedia.org
One main aspect of global workspace theory is that there are specialized modules that operate locally but can be broadcast globally if attention is focused on them. I think LLMs have some sort of workspace, but Anthropic don't claim to find modules, I wouldn't call that GWT.
Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.com/research/glo...
Anthropic has a new paper out, alleging a global workspace in LLMs. The term comes from Global Workspace Theory, a leading theory of consciousness. The method they use to investigate this is a refinement of logit lens, which they call J-lens. www.anthropic.com/research/glo...