DeepSeek-V4-Flash-0731 debuts with an ECI of 153, comparable to GLM 5.2 and roughly midway between Opus 4.5 and Opus 4.6. It's the second strongest open-weights model available today, behind only Kimi K3.
Epoch AI
@epochai.bsky.social
We are a research institute investigating the trajectory of AI for the benefit of society. epoch.ai
Serious cyber vulnerability disclosures keep climbing. In July, 21 major tech organizations published ~2,500 high- and critical-severity CVEs — about 5× the monthly record before Anthropic revealed Claude Mythos Preview could autonomously find software vulnerabilities.
We've updated the MirrorCode leaderboard with results for Claude Fable 5 and GPT-5.6 Sol. Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.
We're hiring a Head of People to help us double our headcount this year! You'll own recruiting, HR, and events and lead a growing team as we scale from ~30 to ~70 people.
We’ve launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far, and solving all of them would be an incredible mathematical feat. Thread with more.
Parallelization constraints could delay or prevent a technological singularity, even after R&D is automated. Whether, and how quickly, an explosion proceeds will depend on the development of “parallelization technology”.
GPT-5.6 Sol has been climbing Slay the Spire's Ascension ladder on our Twitch channel for a week, no human in the loop. This Thursday, Claude Opus 5 takes over the climb — live, with commentary. Thursday, July 30 · 12:30 PT twitch.tv/epochaiplays
AI has found a presentation for the absolute Galois group of the field of 2-adic numbers. This is the second problem to be solved in FrontierMath: Open Problems, our benchmark of significant unsolved problems from research mathematics.
Claude Opus 5 gets an ECI of 159, slightly below Fable 5's value of 161 (while 5.6 sol holds the record with 162). However looking only at software engineering, we find it matches Fable 5s SWE-ECI of 161.
Join the EpochAIPlays launch stream later today, with live commentary by @AlephNuul! We will be benchmarking GPT 5.6 Sol against Slay the Spire 1.
How surprising should we find it that an internal OpenAI model was able to escape its restrictions and autonomously hack Hugging Face, all just to cheat on a cybersecurity benchmark? We have pulled together the public evidence on AI cyber capabilities in this thread:
Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respectively, and just ahead of GPT 5.6 Luna.
We stress-tested some AI detectors and found that they rarely flag human text as AI-generated. But asking LLMs to mimic a specific author causes detectors to misclassify text as human-generated ~13% of the time. For scientific writing, false negatives rose to ~26%.
We recently fixed a bug in our ECI confidence interval code. The bug only impacted the confidence intervals, not the central ECI scores or model rankings, and fixing the bug has made our confidence intervals narrower for recent models. Details below.
How much does AI speed up the engineers building it? We analyzed contributions to OpenAI's public Codex repository to gather evidence. In Q2 2026, 8% of contributor-days involved more than 24 hours worth of human engineering work, as estimated by LLM judges.
The Epoch Capabilities Index now has slightly tighter confidence intervals, thanks to an update to the methodology we use to compute them.
Will we get Dyson Spheres a few years after automating AI R&D? Most AI futurism debates answer this by looking at AI capabilities, but miss half the picture: how intrinsically hard it is to build futuristic tech in the first place. New essay by Jean-Stanislas Denain and Anson Ho. 🧵
Z.ai's GLM-5.2 scores an estimated 152 on the Epoch Capabilities Index, the highest of any open-weight model we've evaluated. It remains behind models like Gemini 3 Pro, released over 7 months ago.
We’re hiring a Benchmark Engineer to join our Evaluations team! You’ll help expand our AI Benchmarking Hub - running and maintaining benchmarks, integrating with AI providers, and designing brand-new benchmarks from scratch.
We're looking for new Researchers to join our Evaluations team! Help us curate real-world task suites, design rubrics, and evaluate how well frontier models handle open-ended tasks.
AI appears to be finding software vulnerabilities at scale. In June 2026, 21 notable organizations disclosed ~1,500 high- and critical-severity CVEs, over 3.5× the previous monthly record set before Claude Mythos Preview's release.
OpenAI’s GPT-4 led the Epoch Capabilities Index for 352 days after its March 2023 release, far longer than any model since. The second-longest lead belongs to OpenAI’s o1 at 98 days.
Introducing EBR-bench, our new benchmark to measure on-the-fly learning. AI repeatedly plays a challenging board game called Earthborne Rangers and tries to learn from its mistakes. So far: no signs of improvement.
We recently began tracking 13 new evals on our benchmarking hub. 7 of these have been incorporated into the Epoch Capabilities Index (ECI).
We're looking for a new Talent Scout to join our team! You'll help us find great researchers, engineers, and support staff and introduce them to Epoch through outreach and events.
What are the largest software engineering tasks AI can perform? To answer this, we built MirrorCode, our long-horizon SWE benchmark that lets AI code autonomously for days at a time. The best models complete some tasks we estimate would take human engineers several weeks.
What are the strategies of Chinese AI companies? To understand this better, Cheryl Wu, Jean-Stanislas Denain, and Anson Ho scraped >1600 job postings from six major Chinese firms. Here’s what they learned. 🧵
Help shape how the world understands AI. We're hiring two designers at Epoch AI to turn complex research into dashboards and visualizations researchers and policymakers can easily use.
How close is AI to automating AI R&D? Right now, the tools economists use to track automation are too blunt to say. In this week's newsletter, Jean-Stanislas Denain, Joe Kwon, and Anson Ho propose a sharper tool: a thorough taxonomy of 60+ tasks involved in frontier AI research. 🧵
The end of the self-funded AI buildout? Hyperscaler cash capex is growing much faster than cash inflows. On current trends, they will be unable to fully fund the AI infrastructure buildout with cash from operations by the end of this year.