Seth Karten

@sethkarten.ai

Autonomous Agents | Research @ Prime Intellect | PhD @ Princeton | Prev: CMU, Waymo | NSF GRFP Fellow https://sethkarten.ai/

Wow, three papers in CoLM 2026... Here I come San Francisco! These papers predicted some early trends in multi-agent safety & economic envs, automatic RL env creation, and PPO for VLMs/LLMs 🧵

Bild

New blog applying Continual Harness to ARC-AGI-3. The heavy test-time learning required by the benchmark pushes agents to form an internal world model of the rules and mechanics that updates with new evidence. Continual Harness scored 20.54%. sethkarten.substack.com/p/continual-...

Continual Harness: An Efficient Self-Improving Agent on ARC-AGI-3

Continual Harness scores 20.54% on ARC-AGI-3 at $774, showing how reset-free self-improving agents can learn hidden game dynamics at test time.

open.substack.com

im very glad to see a rebound this year in the total number of NSF GRFP awards to exceed the most this century. Founding the next generation of American scientists is important to keep growth of the sciences.

Bild

🚨New preprint! LLM teams are being deployed at scale, yet we lack the tools to predict when they’ll succeed, fail, or how to design them. Distributed computing faced the exact same questions and figured out how to answer them. We show those insights apply directly to LLMs 🧵👇

Bild

How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning)

Flyer for The PokeAgent Challenge at NeurIPS 2025. Sunday, Dec 7, 8–10:45 AM PST, Mezzanine Room 15AB, San Diego Convention Center. Two tracks: Track 1 (Battling) features competitive Pokémon battle bots; Track 2 (Speedrunning) features long-horizon RPG gameplay. Tagline: "How do we close the gap between specialist RL models and generalist LLM agents?" Speakers: Seth Karten (Princeton), Aaron Traylor, Minmin Chen (Google DeepMind), Jake Grigsby (UT Austin), Stephanie Milani (NYU/Johns Hopkins), Kiran Vodrahalli (Google DeepMind), Fei Fang (CMU), Yuke Zhu (UT Austin), Chi Jin (Princeton). Sponsored by Google DeepMind.

I’ll be in San Diego at NeurIPS Dec 3-7! DM or email if you want to chat about - building the foundation agents through games - PokeAgent Challenge & PokéChamp - LLM Economist & autonomous business agents

How do we close the gap between specialist RL and generalist LLM agents? We're benchmarking it in Pokémon. Join us at the PokeAgent Challenge competition workshop @ NeurIPS 2025. 📍 Dec 7, 8AM 🎮 Track 1: Competitive Pokémon (game-theoretic reasoning) 🗺️ Track 2: Speedrunning (long-horizon planning)

Flyer for The PokeAgent Challenge at NeurIPS 2025. Sunday, Dec 7, 8–10:45 AM PST, Mezzanine Room 15AB, San Diego Convention Center. Two tracks: Track 1 (Battling) features competitive Pokémon battle bots; Track 2 (Speedrunning) features long-horizon RPG gameplay. Tagline: "How do we close the gap between specialist RL models and generalist LLM agents?" Speakers: Seth Karten (Princeton), Aaron Traylor, Minmin Chen (Google DeepMind), Jake Grigsby (UT Austin), Stephanie Milani (NYU/Johns Hopkins), Kiran Vodrahalli (Google DeepMind), Fei Fang (CMU), Yuke Zhu (UT Austin), Chi Jin (Princeton). Sponsored by Google DeepMind.

Every LLM eval uses Bradley-Terry Elo rankings. Almost none report uncertainty. Should we trust them? Maybe there is something better... 👇 (1/5)

Pokemon is truly the pareto frontier of agent research - The RPG requires an autonomous embodied agentic agent with perception, planning, memory, and control - VGC and Gen 9 OU penalize erroneous actions with fast-paced opponent-modeling in short games (1/3)

You probably aren’t reading enough papers. You probably didn’t cite the 10 closest papers to your work Thus, LLMs probably have a better understanding of where your paper sits in the literature ¯\_(ツ)_/¯

🚨 Hackathon Weekend! 🚨 Jumpstart your PokéAgent Challenge submission ahead of NeurIPS! 📅 Sept 13–14 ✅ Leaderboards reset Sat 10AM EDT 🎙️ Lightning talks in LLMs, RL, and Pokemon 💬 Live Office hours 🏆 $2k in prizes

PokéAgent Challenge @ NeurIPS 2025 Hackathon Weekend Schedule. Saturday, Sept 13th: 10 AM leaderboards reset; 12–1:30 PM livestream talks (overview, Aaron Traylor on Pokémon as an AI Problem, Seth Karten on Pokéchamp, Jake Grigsby on Metamon, plus more). Sunday, Sept 14th: 1–3:30 PM organizer office hours; 11:59 PM top teams earn up to $2k in GCP credits. Sponsored by Google DeepMind and AIJ.