Edoardo Ponti

@edoardo-ponti.bsky.social

Assistant professor in Natural Language Processing at the University of Edinburgh and visiting professor at NVIDIA | A Kleene star shines on the hour of our meeting.

🚀 By *learning* to compress the KV cache in Transformer LLMs, we can generate more tokens for the same compute budget. This unlocks *inference-time hyper-scaling* For the same runtime or memory load, we can boost LLM accuracy by pushing reasoning even further!

Bild

We propose Neurosymbolic Diffusion Models! We find diffusion is especially compelling for neurosymbolic approaches, combining powerful multimodal understanding with symbolic reasoning 🚀 Read more 👇

Sparse attention is one of the most promising strategies to unlock long-context processing and long-generation reasoning in LLMs. We performed the most comprehensive study on training-free sparse attention to date. Here is what we found:

Bild

We created Approximate Likelihood Matching, a principled (and very effective) method for *cross-tokenizer distillation*! With ALM, you can create ensembles of models from different families, convert existing subword-level models to byte-level and a bunch more🧵

Image illustrating that ALM can enable Ensembling, Transfer to Bytes, and general Cross-Tokenizer Distillation.

**Grounded typology**: a new paradigm. Traditionally, linguists posit functions to compare forms in different languages; however, these are aprioristic and partly arbitrary. Instead, we resort to perceptual modalities (like vision) as measurable proxies for function.

Coleman Haley@colemanhaley.bsky.social · 2y ago

NEW PREPRINT! Language is not just a formal system—it connects words to the world. But how do we measure this connection in a cross-linguistic, quantitative way? 🧵 Using multimodal models, we introduce a new approach: groundedness ⬇️