@algoresearch.bsky.social

There's rightly a lot of excitement around Karpathy's autoresearch. We've been studying at ARG for a couple years now: what happens when you put an AI agent in a loop and let it run experiments, evaluate results, and iterate without you. We've built a bunch of benchmarks and tools to measure this. 🧵