Epoch AI

@epochai.bsky.social

We are a research institute investigating the trajectory of AI for the benefit of society. epoch.ai

DeepSeek-V4-Flash-0731 debuts with an ECI of 153, comparable to GLM 5.2 and roughly midway between Opus 4.5 and Opus 4.6. It's the second strongest open-weights model available today, behind only Kimi K3.

Bild

Serious cyber vulnerability disclosures keep climbing. In July, 21 major tech organizations published ~2,500 high- and critical-severity CVEs — about 5× the monthly record before Anthropic revealed Claude Mythos Preview could autonomously find software vulnerabilities.

Bild

We've updated the MirrorCode leaderboard with results for Claude Fable 5 and GPT-5.6 Sol. Claude Fable 5 leads with a 64% solve rate, followed by GPT-5.6 Sol at 20%.

Bild

We're hiring a Head of People to help us double our headcount this year! You'll own recruiting, HR, and events and lead a growing team as we scale from ~30 to ~70 people.

Bild

We’ve launched an expansion of FrontierMath: Open Problems! The benchmark now contains 50 significant, unsolved problems from research mathematics. AI has solved three so far, and solving all of them would be an incredible mathematical feat. Thread with more.

Bild

Parallelization constraints could delay or prevent a technological singularity, even after R&D is automated. Whether, and how quickly, an explosion proceeds will depend on the development of “parallelization technology”.

Bild

AI has found a presentation for the absolute Galois group of the field of 2-adic numbers. This is the second problem to be solved in FrontierMath: Open Problems, our benchmark of significant unsolved problems from research mathematics.

Bild

Claude Opus 5 gets an ECI of 159, slightly below Fable 5's value of 161 (while 5.6 sol holds the record with 162). However looking only at software engineering, we find it matches Fable 5s SWE-ECI of 161.

Bild

How surprising should we find it that an internal OpenAI model was able to escape its restrictions and autonomously hack Hugging Face, all just to cheat on a cybersecurity benchmark? We have pulled together the public evidence on AI cyber capabilities in this thread:

Bild

Moonshot's Kimi K3 scores 156 on the Epoch Capabilities Index (ECI), setting a new open-weights record. This places it between Opus 4.6, and GPT 5.4, which released in February and March 2026 respectively, and just ahead of GPT 5.6 Luna.

Bild

We stress-tested some AI detectors and found that they rarely flag human text as AI-generated. But asking LLMs to mimic a specific author causes detectors to misclassify text as human-generated ~13% of the time. For scientific writing, false negatives rose to ~26%.

Bild

We recently fixed a bug in our ECI confidence interval code. The bug only impacted the confidence intervals, not the central ECI scores or model rankings, and fixing the bug has made our confidence intervals narrower for recent models. Details below.

How much does AI speed up the engineers building it? We analyzed contributions to OpenAI's public Codex repository to gather evidence. In Q2 2026, 8% of contributor-days involved more than 24 hours worth of human engineering work, as estimated by LLM judges.

Bild

Will we get Dyson Spheres a few years after automating AI R&D? Most AI futurism debates answer this by looking at AI capabilities, but miss half the picture: how intrinsically hard it is to build futuristic tech in the first place. New essay by Jean-Stanislas Denain and Anson Ho. 🧵

Bild

Z.ai's GLM-5.2 scores an estimated 152 on the Epoch Capabilities Index, the highest of any open-weight model we've evaluated. It remains behind models like Gemini 3 Pro, released over 7 months ago.

Bild

We’re hiring a Benchmark Engineer to join our Evaluations team! You’ll help expand our AI Benchmarking Hub - running and maintaining benchmarks, integrating with AI providers, and designing brand-new benchmarks from scratch.

Bild

We're looking for new Researchers to join our Evaluations team! Help us curate real-world task suites, design rubrics, and evaluate how well frontier models handle open-ended tasks.

Bild

AI appears to be finding software vulnerabilities at scale. In June 2026, 21 notable organizations disclosed ~1,500 high- and critical-severity CVEs, over 3.5× the previous monthly record set before Claude Mythos Preview's release.

Bild

OpenAI’s GPT-4 led the Epoch Capabilities Index for 352 days after its March 2023 release, far longer than any model since. The second-longest lead belongs to OpenAI’s o1 at 98 days.

Bild

Introducing EBR-bench, our new benchmark to measure on-the-fly learning. AI repeatedly plays a challenging board game called Earthborne Rangers and tries to learn from its mistakes. So far: no signs of improvement.

Bild

We're looking for a new Talent Scout to join our team! You'll help us find great researchers, engineers, and support staff and introduce them to Epoch through outreach and events.

Bild

What are the largest software engineering tasks AI can perform? To answer this, we built MirrorCode, our long-horizon SWE benchmark that lets AI code autonomously for days at a time. The best models complete some tasks we estimate would take human engineers several weeks.

What are the strategies of Chinese AI companies? To understand this better, Cheryl Wu, Jean-Stanislas Denain, and Anson Ho scraped >1600 job postings from six major Chinese firms. Here’s what they learned. 🧵

Bild

Help shape how the world understands AI. We're hiring two designers at Epoch AI to turn complex research into dashboards and visualizations researchers and policymakers can easily use.

Bild

How close is AI to automating AI R&D? Right now, the tools economists use to track automation are too blunt to say. In this week's newsletter, Jean-Stanislas Denain, Joe Kwon, and Anson Ho propose a sharper tool: a thorough taxonomy of 60+ tasks involved in frontier AI research. 🧵

Bild

The end of the self-funded AI buildout? Hyperscaler cash capex is growing much faster than cash inflows. On current trends, they will be unable to fully fund the AI infrastructure buildout with cash from operations by the end of this year.

Bild