Jay Alammar

@jayalammar.bsky.social

Writer http://jalammar.github.io. O'Reilly Author http://LLM-book.com. LLM Builder Cohere.com.

Inside NeurIPS 2025: The Year’s AI Research, Mapped New blog post! NeurIPS 2025 papers are out—and it’s a lot to take in. This visualization lets you explore the entire research landscape interactively, with clusters and @cohere.com LLM-generated explanations that make it easier to grasp.

The legendary John Carmack at #upperbound: - Current AI focus is RL (with Richard Sutton) solving Atari games - Thinking in line with the Alberta Plan. - It was a misstep to start working too low-level (e.g., at the cuda level). I kept stepping up the stack chain until now in pytorch

Bild

I'm really excited for this year's PyData London conference - there are some awesome talks on the schedule and I'm excited to hear the keynote speakers @jayalammar.bsky.social, Tony Wears, & Leanne Fitzpatrick #pydata #datascience

PyData London@pydatalondon.bsky.social · last yr.

Unleash your inner data aficionado at PyData London 2025, 6-8 June at Convene Sancroft, St. Paul’s! We have 3 top flight keynotes lined up for you this year from @jayalammar.bsky.social, Leanne Kim Fitzpatrick and Tony Mears. Just 17 days left. Book your tickets now! pydata.org/london2025

Advertisement for PyData London 2025 conference.

Headline: Meet your keynote speakers

- Jay Alammar
- Tony Mears
- Leanne Fitzpatrick

Book your tickets
https://pydata.org/london2025

We've just released MMTEB, our multilingual upgrade to the MTEB Embedding Benchmark! It's a huge collaboration between 56 universities, labs, and organizations, resulting in a massive benchmark of 1000+ languages, 500+ tasks, and a dozen+ domains. Details in 🧵

Bild

One of my grand interpretability goals is to improve human scientific understanding by analyzing scientific discovery models, but this is the most convincing case yet that we CAN learn from model interpretation: Chess grandmasters learned new play concepts from AlphaZero's internal representations.

Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero

Artificial Intelligence (AI) systems have made remarkable progress, attaining super-human performance across various domains. This presents us with an opportunity to further human knowledge and improv...

arxiv.org

The newest extremely strong embedding model based on ModernBERT-base is out: `cde-small-v2`. Both faster and stronger than its predecessor, this one tops the MTEB leaderboard for its tiny size! Details in 🧵

Bild

Floored that the repo for Hands-On Large Language Models is now at 3.6k Github stars! And excited that professors are starting to use the book to teach LLM courses. Reach out to us if we can be of assistance! And if you've liked the book, leave us a review on Amazon or Goodreads!

Bild

🚨 LLMs can learn to reason from procedural knowledge in pretraining data! 🚨 I particularly enjoy research where the evidence contradicts our initial hypothesis. If you're interested in LLM reasoning, check out the 60+ pages of in-depth work at arxiv.org/abs/2411.12580

Laura@lauraruis.bsky.social · 2y ago

How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️