Ben Lipkin

@benlipkin.bsky.social

researcher. prev: phd @ mit, intern @ apple https://benlipkin.github.io/

Humans largely learn language through speech. In contrast, most LLMs learn from pre-tokenized text. In our #Interspeech2025 paper, we introduce AuriStream: a simple, causal model that learns phoneme, word & semantic information from speech. Poster P6, tomorrow (Aug 19) at 1:30 pm, Foyer 2.2!

Many LM applications may be formulated as text generation conditional on some (Boolean) constraint. Generate a… - Python program that passes a test suite. - PDDL plan that satisfies a goal. - CoT trajectory that yields a positive reward. The list goes on… How can we efficiently satisfy these? 🧵👇

I might be able to hire a postdoc for this fall in computational linguistics at UT Austin. Topics in the general LLM + cognitive space (particularly reasoning, chain of thought, LLMs + code) and LLM + linguistic space. If this could be of interest, feel free to get in touch!