Marco Cuturi

@marcocuturi.bsky.social

machine learning researcher @ Apple machine learning research

Small Language Models (SLMs) don’t have the capacity to remember everything in their training data. Which tokens should they learn to predict, and when should they ask for help? We tackle this question in our new preprint. You can check it out on arxiv: arxiv.org/abs/2602.12005 🧵1/7

Bild

With other folks at 🍏, @brunokm.bsky.social has worked on a complete(d) parameterisation for NNs that can *transfer* locally tuned hyperparameters: tune optimizers' parameters (e.g. LR) *per module/depth* using an evolutionary search on small models → they transfer perf. gains to much larger models

Bruno Mlodozeniec@brunokm.bsky.social · 7mo ago

In our new work — Complete(d)P — we try to answer 3 questions about hyperparameter (HP) scaling: ● How to transfer across model size, tokens&batch-size?→ Complete(d)P ● Do per-module HPs matter? ✔️2x speed-ups possible ● Do they transfer to larger scale? ✔️ With the right parameterisation

my 2 cents on the ICLR drama: It's been years that the system has been under attack. But it's also been years that we hear, year after year, that there is no way to enforce protection mechanisms (e.g. deny lists for dishonest authors or reviewers etc..) for legal reasons.

📢 We’re looking for a researcher in in cogsci, neuroscience, linguistics, or related disciplines to work with us at Apple Machine Learning Research! We're hiring for a one-year interdisciplinary AIML Resident to work on understanding reasoning and decision making in LLMs. 🧵

🚀 Excited to share LinEAS, our new activation steering method accepted at NeurIPS 2025! It approximates optimal transport maps e2e to precisely guide 🧭 activations achieving finer control 🎚️ with ✨ less than 32 ✨ prompts! 💻https://github.com/apple/ml-lineas 📄https://arxiv.org/abs/2503.10679

Our two phenomenal interns, Alireza Mousavi-Hosseini and Stephen Zhang @syz.bsky.social have been cooking some really cool work with Michal Klein and me over the summer. Relying on optimal transport couplings (to pick noise and data pairs) should, in principle, be helpful to guide flow matching 🧵

Bild

So pleased and proud to share with you what our team has been up to, on an ambitious journey to build a video foundation model for scientific domains ! ✨ 🚀 🎞️ 🧪 #ICCV2025 #AI4Science

@hassony2.bsky.social · last yr.

Thrilled to share our latest work on SciVid, to appear at #ICCV2025! 🎉 SciVid offers cross-domain evaluation of video models in scientific applications, including medical CV, animal behavior, & weather forecasting 🧪🌍📽️🪰🐭🫀🌦️ 📝 Check out our paper: arxiv.org/abs/2507.03578 [1/4]🧵

NEW PAPER ALERT: Recent studies have shown that LLMs often lack robustness to distribution shifts in their reasoning. Our paper proposes a new method, AbstRaL, to augment LLMs’ reasoning robustness, by promoting their abstract thinking with granular reinforcement learning.

Bild

The trader made millions. Why was this unusual? For a few reasons: - Firstly new opening volume on chain, minutes after market open - IVR of +80 on $QQQ, with iv percentile of medium. - happened at once, at ask -otm - market was bearish, across the board Unusual. Come learn: unusualwhales.com

Bild