Dhruv Batra

@dhruvbatra.bsky.social

Co-founder & Chief Scientist at Yutori. Prev: Senior Director leading FAIR Embodied AI at Meta, and Professor at Georgia Tech.

We eval'd Muse Spark 1.1 on Online-Mind2Web — a computer-use / browser-use benchmark • Better than Opus 4.8 • Slightly worse than GPT 5.4 (but possibly not a statistically significant difference)

Bild

𝗡𝗮𝘃𝗶𝗴𝗮𝘁𝗼𝗿 𝗻𝟭.𝟱 “𝘀𝗼𝗹𝘃𝗲𝗱” 𝗢𝗻𝗹𝗶𝗻𝗲 𝗠𝗶𝗻𝗱𝟮𝗪𝗲𝗯: 𝟵𝟳.𝟯% 𝘀𝘂𝗰𝗰𝗲𝘀𝘀 𝗿𝗮𝘁𝗲. While some teams self-report, this result is independently evaluated and verified by OSU NLP Group and Careerflow Human Data Labs.

Bild

𝐈𝐧𝐭𝐫𝐨𝐝𝐮𝐜𝐢𝐧𝐠 𝐍𝐚𝐯𝐢𝐠𝐚𝐭𝐨𝐫 𝐧𝟏.𝟓 The most capable computer-use model for the web. Pareto-domination: accuracy, latency, cost • SoTA across all benchmarks • +5-10% over GPT 5.5, Opus 4.7, n1 • +25% over Gemini • 2x faster, significantly cheaper

Bild

I gave Claude Code & Codex a video of @yutori_ai Navigator logo spinning and asked for code to regenerate it. Opus 4.7 max (left) vs GPT 5.4 xhigh (right) GPT 5.4 clearly better. Ground-truth / OG video in 🧵

Two updates from Yutori: 1. We benchmarked GPT 5.4 on browser-use tasks • Matches/slightly-outperforms Opus 4.6 (+0.3%) • Big jump over previous OpenAI CUAs 2. Latest version of n1 • Outperforms GPT 5.4 and Opus 4.6 (+3%) • 2.5x faster, 4-5x cheaper.

Bild

Most recent checkpoint of n1 vs Opus 4.6! On Navi-Bench and Westworld browser automation benchmarks: - Same accuracy - n1 is 2.5x faster - n1 is 5.6x cheaper Try it out via the Yutori API.

Bild

The bitter lesson for web agents The last 1 year has taught us a new bitter lesson that we think others are not yet grokking. Agents that *look at the web like humans* (screenshots of sites) navigate and generalize better than agents that read code (HTML, DOM).

Bild

VQA challenge series won the Mark Everingham prize at #ICCV2025 for stimulating a new strand of vision-and-language research. It's extra special because ICCV25 marks the 10-year anniversary of the VQA paper. When we started, the idea of answering any question about any image seemed outlandish.

BildBildBild

The problem with “AI slop” isn’t the AI — it’s the slop. People act like AI is the issue, when it’s actually part of the fix. If we're honest: most of what we make, most of the time, is slop by our own standards. That’s the generator–discriminator gap in creative work that Ira Glass talks about.

I started something new last year with a wonderful group of people. We showed a demo in Jan. Today, we’re telling our story — show before you talk! 𝘞𝘦 𝘢𝘳𝘦 𝘳𝘦-𝘪𝘮𝘢𝘨𝘪𝘯𝘪𝘯𝘨 𝘩𝘰𝘸 𝘱𝘦𝘰𝘱𝘭𝘦 𝘪𝘯𝘵𝘦𝘳𝘢𝘤𝘵 𝘸𝘪𝘵𝘩 𝘵𝘩𝘦 𝘸𝘦𝘣 — one of humanity’s greatest inventions and a a mess overdue for an overhaul. yutori.com

Bild

Brilliant talk by Ilya, but he's wrong on one point. We are NOT running out of data. We are running out of human-written text. We have more videos than we know what to do with. We just haven't solved pre-training in vision. Just go out and sense the world. Data is easy.

Bild