How are the academics feeling about this? Does it even change anything for profs?
Seth Karten
@sethkarten.ai
Autonomous Agents | Research @ Prime Intellect | PhD @ Princeton | Prev: CMU, Waymo | NSF GRFP Fellow https://sethkarten.ai/
Great to see Continual Harness acknowledged in Schmidhuber’s latest survey paper
Going through my backlog and realizing I forgot to announce 1 paper and never put another on arXiv. Expect 2 blog posts soon
Wow, three papers in CoLM 2026... Here I come San Francisco! These papers predicted some early trends in multi-agent safety & economic envs, automatic RL env creation, and PPO for VLMs/LLMs 🧵
New blog applying Continual Harness to ARC-AGI-3. The heavy test-time learning required by the benchmark pushes agents to form an internal world model of the rules and mechanics that updates with new evidence. Continual Harness scored 20.54%. sethkarten.substack.com/p/continual-...
Continual Harness: An Efficient Self-Improving Agent on ARC-AGI-3
Continual Harness scores 20.54% on ARC-AGI-3 at $774, showing how reset-free self-improving agents can learn hidden game dynamics at test time.
open.substack.com
Just know that my reviewers will be thoroughly reviewed for their strengths and weaknesses. Score and confidence included.
New paper alert: Continual Harness: Online Adaptation for Self-Improving Foundation Agents Paper (arXiv). arxiv.org/abs/2605.09998 Article (Substack). sethkarten.substack.com/p/gemini-pla... Project page (video demos). sethkarten.ai/continual-ha...
Announcing some work tomorrow. Will be cool and probably involving pokemon
im very glad to see a rebound this year in the total number of NSF GRFP awards to exceed the most this century. Founding the next generation of American scientists is important to keep growth of the sciences.
New meta seems to be arxiving a rough draft so that you can claim the terminology first and claim to be first
Everyone wants to own their own data but no one wants to own their own data center
🚨New preprint! LLM teams are being deployed at scale, yet we lack the tools to predict when they’ll succeed, fail, or how to design them. Distributed computing faced the exact same questions and figured out how to answer them. We show those insights apply directly to LLMs 🧵👇
I think I accidentally stumbled upon engagement baiting from first principles Ill stay on bluesky as long as the 10 accounts I like to see still post here
I think I might leave bluesky tbh
How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning)
I’ll be in San Diego at NeurIPS Dec 3-7! DM or email if you want to chat about - building the foundation agents through games - PokeAgent Challenge & PokéChamp - LLM Economist & autonomous business agents
How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning)
Every LLM eval uses Bradley-Terry Elo rankings. Almost none report uncertainty. Should we trust them? Maybe there is something better... 👇 (1/5)
Pokemon is truly the pareto frontier of agent research - The RPG requires an autonomous embodied agentic agent with perception, planning, memory, and control - VGC and Gen 9 OU penalize erroneous actions with fast-paced opponent-modeling in short games (1/3)
Trying to get a post ready but bluesky won’t let me post on desktop!!! If you want users here you need a user experience!!!
You probably aren’t reading enough papers. You probably didn’t cite the 10 closest papers to your work Thus, LLMs probably have a better understanding of where your paper sits in the literature ¯\_(ツ)_/¯
The most interesting papers arent being published at the “prestigious” venues anymore. Where are you publishing and what do you work on?
🚨 Hackathon Weekend! 🚨 Jumpstart your PokéAgent Challenge submission ahead of NeurIPS! 📅 Sept 13–14 ✅ Leaderboards reset Sat 10AM EDT 🎙️ Lightning talks in LLMs, RL, and Pokemon 💬 Live Office hours 🏆 $2k in prizes
The NeurIPS 2025 PokéAgent Challenge is offering compute credits, courtesy of our sponsor Google DeepMind, to help you train bigger models & run more experiments. 📌 To apply: 1️⃣ Make a submission to Track 1 or 2 at pokeagent.github.io 2️⃣ Fill out the compute credit form on the site
PokéAgent Challenge - NeurIPS 2025
pokeagent.github.io