Sasha Rush

@srushnlp.bsky.social

Professor, Programmer in NYC. Cornell, Hugging Face 🤗

For reasons, I find myself thinking a lot about the history of US/USSR Cold War science, particularly in applied math. Does anyone have a recommendation for a good book on this topic?

We outperform Llama 70B with Llama 3B on hard math by scaling test-time compute 🔥 How? By combining step-wise reward models with tree search algorithms :) We're open sourcing the full recipe and sharing a detailed blog post 👇

Bild

As coding LLMs get faster at inference, iterating verification-in-the-loop tests becomes the bottleneck for coding agents. Probably need quite different programming systems for these settings, or even things like "batchable" runtimes, whatever that means.

I wanted to make my first post about a project close to my heart. Linear algebra is an underappreciated foundation for machine learning. Our new framework CoLA (Compositional Linear Algebra) exploits algebraic structure arising from modelling assumptions for significant computational savings! 1/4

Bild

Is there a community that writes RL-first programming languages? Something like (Num)Pyro that takes seriously the idea of separating the policy specification from the learning process.

ICYMI, @srushnlp.bsky.social recently gave a nice talk speculating about the methods/data used to train OpenAI's o1 model. The key idea seems to be scaling up chain-of-thought (CoT) generation using auxiliary verifier models that can give feedback on the correctness of the generation. 1/N

Sasha Rush@srushnlp.bsky.social · 2y ago

Talk: Speculations on Test-Time Scaling A tutorial on the technical aspects behind OpenAI's o1 and open research questions in this space. youtu.be/6PEJ96k1kiw Slides+bibliography: github.com/srush/awesom...

Rare personal tweet: Subletting our furnished apartment in Brooklyn for the spring at a significant discount. It's quite nice and in a fun location. under price. Email me know if you are interested, I will send pictures.

I've spent the last two years scouring all available resources on RLHF specifically and post training broadly. Today, with the help of a totally cracked team, we bring you the fruits of that labor — Tülu 3, an entirely open frontier model post training recipe. We beat Llama 3.1 Instruct. Thread.

BildBild

Discrete diffusion has become a very hot topic again this year. Dozens of interesting ICLR submissions and some exciting attempts at scaling. Here's a bibliography on the topic from the Kuleshov group (my open office neighbors). github.com/kuleshov-gro...

GitHub - kuleshov-group/awesome-discrete-diffusion-models: A curated list for awesome discrete diffusion models resources.

A curated list for awesome discrete diffusion models resources. - kuleshov-group/awesome-discrete-diffusion-models

github.com

hard for me to get engagement on this platform but it’s ok I’m just waiting for all the AI researchers to show up so I can go mega viral saying shit like “who called it Xavier Initialization and not Weighting For Glorot”

I keep reading interviews about how impressively smart and accomplished the AlphaProof team is. But I would be way more impressed if they were just like "we're super bad at math, who cares".