Michael Kirchhof

@mkirchhof.bsky.social

RS on uncertainty quantification for agents at Apple

4 ICLR papers 🥳 There’s an insightful story between them: If you sample LLMs multiple times, they are calibrated, even on higher levels [1], but they cannot talk about this uncertainty in a single prompt [2], so you have to help them out to gather information Bayes-optimally [3]

Bild

In our new work — Complete(d)P — we try to answer 3 questions about hyperparameter (HP) scaling: ● How to transfer across model size, tokens&batch-size?→ Complete(d)P ● Do per-module HPs matter? ✔️2x speed-ups possible ● Do they transfer to larger scale? ✔️ With the right parameterisation

Bild

Our research team is hiring PhD interns 🍏 Spend your next summer in Paris and explore the next frontiers of LLMs for uncertainty quantification, calibration, RL and post-training, and Bayesian experimental design. Details & Application ➡️ jobs.apple.com/en-my/detail...

Internship - Machine Learning Research on Uncertainty - Jobs at Apple (MY)

Apply for a Internship - Machine Learning Research on Uncertainty job at Apple. Read about the role and find out if it’s right for you.

jobs.apple.com

📢 We’re looking for a researcher in in cogsci, neuroscience, linguistics, or related disciplines to work with us at Apple Machine Learning Research! We're hiring for a one-year interdisciplinary AIML Resident to work on understanding reasoning and decision making in LLMs. 🧵

Our two phenomenal interns, Alireza Mousavi-Hosseini and Stephen Zhang @syz.bsky.social have been cooking some really cool work with Michal Klein and me over the summer. Relying on optimal transport couplings (to pick noise and data pairs) should, in principle, be helpful to guide flow matching 🧵

Bild

NEW PAPER ALERT: Recent studies have shown that LLMs often lack robustness to distribution shifts in their reasoning. Our paper proposes a new method, AbstRaL, to augment LLMs’ reasoning robustness, by promoting their abstract thinking with granular reinforcement learning.

Bild

I‘ll talk today about our latest research on uncertainty quantification at Apple (papers are 2 weeks old) and what I see as the future for UQ in vision and LLMs. See you at 102B, 4:30pm! PS: Lmk if you wanna chat :)

Bild

🚨 Apple Machine Learning Research Internship opportunity! My colleagues in Apple MLR are looking for a PhD research intern with a strong interest in reinforcement learning/post-training for LLMs. If interested, apply by sending an email to Etai Littwin (elittwin at apple dot com)

Wow, OpenAI's o1 has a whopping 93% ECE on Humanity's Last Exam. So if you just prompt o1 to tell you how sure it is about its answer, it will basically produce gibberish. And that's how most users will ask for uncertainties. We have work to do!

Results of state-of-the-art LLMs on Humanity's Last Exam are surprisingly bad, especially their uncertainties.