Atlas Wang

@atlaswang.bsky.social

https://www.vita-group.space/ 👨‍🏫 UT Austin ML Professor (on leave) https://www.xtxmarkets.com/ 🏦 XTX Markets Research Director (NYC AI Lab) Superpower is trying everything 🪅 Newest focus: training next-generation super intelligence - Preview above 👶

Of all my students’ career achievements, this one makes me proudest: because really, who needs yet another AI professor or scientist alumni when you can boast a professional cat-breeder student? www.bestbritishcats.com To my Bay Area friends: BUY one! Maybe try using my name to get a discount :)

Snow Moon British | 🌴 CA, 🇺🇸 | Affectionate British Kittens for You

Snow Moon is a trusted British cattery located in San Jose, CA, specializing in healthy, champion-quality British Shorthair and Longhair kittens. Ethical breeder. Breeding cats are genetically tested,...

bestbritishcats.com

@talkachman.bsky.social caught it faster before I did... but here we go! 🚀 New pre-print dropped lnkd.in/gmQ9hBdf Its first draft won the #DARPA Disruptive Idea award at #NeuS2025. Now my student Peihao Wang has sharpened the theory even more. I’m so excited what it means for AI reasoning 🧩🤖 (1/n)

LinkedIn

This link will take you to a page that’s not on LinkedIn

lnkd.in

@talkachman.bsky.social · last yr.

@atlaswang.bsky.social knocking out of the park again! half a page in and this is a phenomenal read. So much to wrap my head around www.arxiv.org/abs/2506.21797

Check out #PanEcho, now in @jama.com ✅ open weights ✅ multiview multitask reporting ✅ international validation ✅ works with POCUS Amazing effort @cardslab.bsky.social @giholste.bsky.social @rohankhera.bsky.social + @atlaswang.bsky.social, Márton Tokodi & Attila Kovács #CardioSky #EchoSky

JAMA@jama.com · last yr.

An #AI system that automatically interprets echocardiograms maintained high accuracy across geography and time from complete and limited studies. https://ja.ma/4jYWBgI

Figure 4. Task-Specific ViewRelevance

Making a new website of my research group, and did a visualization of all our papers from 2018 to present: clustered into 10 topics. One can clearly see how this group evolves its own tastes! … and deeper in my heart: long live optimization!! ❤️

Bild

🚀 Thrilled to announce SPIN-Bench!🚀 We all love seeing how smart LLMs can be-solving complex math, crafting beautiful text, and coding effortlessly. But how well do they handle real-world strategic complexity, cooperation, & social negotiation? Can they play well when things get tricky? Not quite!

Bild

Just had a meal that gifted me two rare treasures: 1️⃣ Meeting someone infinitely wiser than me. 2️⃣ They weren’t cold or mean—just gently showed me where I could grow. "To learn truth at dawn, I’d die content by dusk." ✨ Humility tastes better with kindness. #Gratitude #LifeLessons

Good news? SPIN-Bench pinpoints exactly where these models fall short—and illuminates exciting research directions to smarter, socially savvy AI For some extra fun (and detailed, interactive trajectory visualizations!), visit project page: 👉 spinbench.github.io full paper: arxiv.org/pdf/2503.12349

SPIN-Bench: How Well Do Large Language Models Reason Strategically and Socially?

How Well Do Large Language Models Reason Strategically and Socially?

spinbench.github.io

🌍 Diplomacy – The ultimate test! Models had to negotiate, forge alliances, and occasionally backstab. Result? Even the best LLMs floundered, struggling to juggle complex social interactions and strategic depth.

Bild

🎴 Cooperative Games (Hanabi): Coordination among teammates dramatically challenges models, causing their performance to dip sharply as complexity ramps up. Turns out, keeping track of your teammates’ intentions isn’t an easy task—even for GPTs!

Bild

Results fascinatingly reveal: 🔹 Classic Planning: LLMs ace simpler puzzles but struggle badly as complexity grows—losing track at longer-term decisions 🔎 Competitive Games: Top chess engines swept every LLM clean. Even simple tactical awareness quickly fades when facing deeper strategic branches

Bild