Houjun Liu
@jemoka.com
NLP & POMDPs; CS@Stanford; gradient descent enthusiast www: jemoka.com ac: nlp.stanford.edu/~houjun/
LMAO, openreview down point 9pm UCT when ICLR is supposed to be releasing. Coincidence?
Introducing 𝘁𝗵𝗼𝘂𝗴𝗵𝘁𝗯𝘂𝗯𝗯𝗹𝗲𝘀: a *fully unsupervised* LM for input-adaptive parallel latent reasoning ✅ Learn yourself a reasoning model with normal pretraining ✅ Better perplexity compared to fixed thinking tokens No fancy loss, no chain of thought labels 🚀
Introducing 𝘁𝗵𝗼𝘂𝗴𝗵𝘁𝗯𝘂𝗯𝗯𝗹𝗲𝘀: a *fully unsupervised* LM for input-adaptive parallel latent reasoning ✅ Learn yourself a reasoning model with normal pretraining ✅ Better perplexity compared to fixed thinking tokens No fancy loss, no chain of thought labels 🚀
New Paper Day! For EMNLP findings—in LM red-teaming, we show you have to optimize for **both** perplexity and toxicity for high-probability, hard to filter, and natural attacks!
New Paper Day! For EMNLP findings—in LM red-teaming, we show you have to optimize for **both** perplexity and toxicity for high-probability, hard to filter, and natural attacks!
The list of accepted papers at #FOCS2025 is up! focs.computer.org/2025/accepte...
Accepted Papers – FOCS 2025
focs.computer.org
You're not too dumb for Haskell, you just need a reason to practice. :)
Just published in JOSS: 'Turftopic: Topic Modelling with Contextual Representations from Sentence Transformers' https://doi.org/10.21105/joss.08183
We’re proud to announce three new tenure-track assistant professors joining TTIC in Fall 2026: Yossi Gandelsman, Will Merrill, and Nick Tomlin (@nickatomlin.bsky.social). Meet them here: buff.ly/JH1DFtT
New paper on the generalization of Flow Matching www.arxiv.org/abs/2506.03719 🤯 Why does flow matching generalize? Did you know that the flow matching target you're trying to learn *can only generate training points*? w @quentinbertrand.bsky.social @annegnx.bsky.social @remiemonet.bsky.social 👇👇👇
New Paper Day! For ACL 2025 Findings: You should **drop dropout** when you are training your LMs AND MLMs!
New Paper Day! For ACL 2025 Findings: You should **drop dropout** when you are training your LMs AND MLMs!
I'm excited to announce that I’ll be joining the Computer Science department at Johns Hopkins as an Assistant Professor this Fall! I’ll be working on large language models, computational social science, and AI & society—and will be recruiting PhD students. Apply to work with me!
looks like deepmind has been busy deepmind.google/models/veo/ deepmind.google/models/gemin...
Jax has a debugger now! This changes everything: docs.jax.dev/en/latest/de...
Compiled prints and breakpoints — JAX documentation
docs.jax.dev
Inspiring celebration of David Attenborough from Kate Winslet and many others - and a much-needed reminder to stand with science www.theguardian.com/tv-and-radio...
Happy birthday, David Attenborough! 99 ways he has inspired us, by Barack Obama, Billie Eilish, Morgan Freeman – and many more
This week the presenter turns 99. To celebrate, we asked 99 nature lovers – including Margaret Atwood, Jane Fonda, Bono, Kate Winslet and Michael Palin – how he has helped us see the world with fresh ...
theguardian.com
Is it just me or is the latest style guide of ChatGPT's IFT is like terribly sarcastic? I don't need 👉 finger guns after every single message.
I'm to this day still confused about why people are so hyped about softmax-bottlenecked "deep research" approaches; idk about you but I don't usually need to compose a 10 page treatise on first order logic to decide whether or not an if statement is backwards....
cs theory seems like an outlier within theoretical sciences where we are living in the era of Euler, Gauss, etc. for math with substantial results being continuously developed live as a part of frontier research. what a time to be alive.
New paper: Simulating Time With Square-Root Space people.csail.mit.edu/rrw/time-vs-... It's still hard for me to believe it myself, but I seem to have shown that TIME[t] is contained in SPACE[sqrt{t log t}]. To appear in STOC. Comments are very welcome!
the world if I could spell "causal interventions" correctly on the first try
Ever dreamed of AI agents learning through interacting with the open world unsupervisedly? Our latest preprint introduces NNetNav-Live which collects training data through exploration on real websites and hindsight labeling, which produces a SOTA OSS agent.
LM agents today primarily aim to automate tasks. Can we turn them into collaborative teammates? 🤖➕👤 Introducing Collaborative Gym (Co-Gym), a framework for enabling & evaluating human-agent collaboration! I now get used to agents proactively seeking confirmations or my deep thinking.(🧵 with video)
shout out to @sfcompute.bsky.social for saving my last-minute compute ass
Happy New Year everyone! Jim and I just put up our January 2025 release of Speech and Language Processing! Check it out here: web.stanford.edu/~jurafsky/sl...
Speech and Language Processing
Speech and Language Processing
web.stanford.edu
I feel like the instagram onboarding is at this point the primary blocker of me getting an instagram. They *insist* that I’m {fake, spam, using an open proxy} literally no matter what I try or where I try it. So be it, better that way probably too :/ But I do want to post pictoors… maybe Pinterest?
Come visit our poster "MoEUT: Mixture-of-Experts Universal Transformers" on Friday at 4:30 in East Exhibit Hall A-C #1907 on #NeurIPS2024. With Kazuki Irie, Jürgen Schmidhuber, Christopher Potts and @chrmanning.bsky.social.