A look at what’s next in AI, plus new research in multimodal reasoning, long-horizon robotics, scalable self-supervised learning, GPU optimization with AI, and interpretable LLM reasoning. msft.it/6012tfazi
Lester Mackey
@lestermackey.bsky.social
Machine learning researcher at Microsoft Research. Adjunct professor at Stanford.
Eric Zimmermann, Harley Wiltzer, Justin Szeto, David Alvarez-Melis, Lester Mackey: KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning https://arxiv.org/abs/2512.19605 https://arxiv.org/pdf/2512.19605 https://arxiv.org/html/2512.19605
I’m happy to share our new paper, “KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning”: arxiv.org/abs/2512.19605 This work builds on the LeJEPA framework by reinterpreting its regularization objective through the lens of kernel discrepancies.
If you're a PhD student interested in interning with me or one of my amazing colleagues at Microsoft Research New England @msftresearch.bsky.social this summer, please apply here jobs.careers.microsoft.com/global/en/jo... If you'd like to work with me, please include my name in your cover letter!
Search Jobs | Microsoft Careers
jobs.careers.microsoft.com
Microsoft Research New York City (www.microsoft.com/en-us/resear...) is seeking applicants for multiple Postdoctoral Researcher positions in ML/AI! These are positions for up to 2 years, starting in July 2026. Application deadline: October 22, 2025
Microsoft Research Lab - New York City - Microsoft Research
Apply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML.
microsoft.com
MSR NYC is hiring spring and summer interns in AI/ML/RL! Apply here: jobs.careers.microsoft.com/global/en/jo...
Microsoft Research Lab - New York City - Microsoft Research
Apply for a research position at Microsoft Research New York & collaborate with academia to advance economics research, prediction markets & ML.
microsoft.com
Very excited to share our preprint: Self-Speculative Masked Diffusions We speed up sampling of masked diffusion models by ~2x by using speculative sampling and a hybrid non-causal / causal transformer arxiv.org/abs/2510.03929 w/ @vdebortoli.bsky.social, Jiaxin Shi, @arnauddoucet.bsky.social
You can read more in our post at www.jurabio.com/blog/leavs; preprint forthcoming. @jura.bsky.social @eliweinstein.bsky.social @mgollub.bsky.social @highvariance.bsky.social
LeaVS: Accelerating learning for biological AI — JURA Bio, Inc.
A fundamental lesson of modern AI is that scale is essential: training bigger models on bigger datasets unlocks new capabilities. A fundamental lesson of AI engineering is that scaling up isn't trivia...
jurabio.com
The Microsoft Research Undergraduate Internship Program offers 12-week internships in our Redmond, NYC, or New England labs for rising juniors and seniors who are passionate about technology. Apply by October 6: msft.it/6015scgSJ
If you're an undergraduate interested in interning with me or one of my amazing colleagues at Microsoft Research New England this summer, please apply here: msft.it/6015scgSJ
Undergraduate Research Internship – Computing - Microsoft Research
Accepting applications for 12-week summer research internships for juniors & senior undergrads w/ demonstrated leadership in diversity.
msft.it
The Microsoft Research Undergraduate Internship Program offers 12-week internships in our Redmond, NYC, or New England labs for rising juniors and seniors who are passionate about technology. Apply by October 6: msft.it/6015scgSJ
Tomorrow we're excited to host @sarahalamdari.bsky.social at Chalmers for the AI4Science seminar and hear about generative models for protein design! Talk at 3pm CEST. 🤩 For more info, including details on how to join virtually, please see psolsson.github.io/AI4ScienceSe... @smnlssn.bsky.social
We may have the chance to hire an outstanding researcher 3+ years post PhD to join Tarleton Gillespie, Mary Gray and me in Cambridge MA bringing critical sociotechnical perspectives to bear on new technologies. jobs.careers.microsoft.com/global/en/jo...
Search Jobs | Microsoft Careers
jobs.careers.microsoft.com
In 1965, Margaret Dayhoff published the Atlas of Protein Sequence and Structure, which collated the 65 proteins whose amino acid sequences were then known. Inspired by that Atlas, today we are releasing the Dayhoff Atlas of protein sequence data and protein language models.
So you want to skip our thinning proofs—but you’d still like our out-of-the-box attention speedups? I’ll be presenting the Thinformer at two ICML workshop posters tomorrow! Catch me at Es-FoMo (1-2:30, East hall A) and at LCFM (10:45-11:30 & 3:30-4:30, West 202-204)
Your data is low-rank, so stop wasting compute! In our new paper on low-rank thinning, we share one weird trick to speed up Transformer inference, SGD training, and hypothesis testing at scale. Come by ICML poster W-1012 Tuesday at 4:30!
Jikai Jin, Lester Mackey, Vasilis Syrgkanis: It's Hard to Be Normal: The Impact of Noise on Structure-agnostic Estimation https://arxiv.org/abs/2507.02275 https://arxiv.org/pdf/2507.02275 https://arxiv.org/html/2507.02275
Off to ICML next week? Check out my student Annabelle’s paper in collaboration with @lestermackey.bsky.social and colleagues on low-rank thinning! New theory, dataset compression, efficient attention and more: arxiv.org/abs/2502.12063
Low-Rank Thinning
The goal in thinning is to summarize a dataset using a small set of representative points. Remarkably, sub-Gaussian thinning algorithms like Kernel Halving and Compress can match the quality of unifor...
arxiv.org
Your data is low-rank, so stop wasting compute! In our new paper on low-rank thinning, we share one weird trick to speed up Transformer inference, SGD training, and hypothesis testing at scale. Come by ICML poster W-1012 Tuesday at 4:30!
New guarantees for approximating attention, accelerating SGD, and testing sample quality in near-linear time
Jikai Jin, Lester Mackey, Vasilis Syrgkanis It's Hard to Be Normal: The Impact of Noise on Structure-agnostic Estimation https://arxiv.org/abs/2507.02275
NeurIPS is seeking additional ethics reviewers this year. If you are able and willing to participate in the review process, please sign up at the form in the link: neurips.cc/Conferences/... Please share this call with your colleagues!
2025 Call For Ethics Reviewers
If you are able and willing to participate in the review process, please sign up at this form. Feel free to share this call with your colleagues.
neurips.cc
Shunichi Amari has been awarded the 40th (2025) Kyoto Prize in recognition of his pioneering research in the fields of artificial neural networks, machine learning, and information geometry www.riken.jp/pr/news/2025...
甘利 俊一 栄誉研究員が「京都賞」を受賞
甘利 俊一栄誉研究員(本務:帝京大学 先端総合研究機構 特任教授)は、人工ニューラルネットワーク、機械学習、情報幾何学分野での先駆的な研究が評価され、第40回(2025)京都賞(先端技術部門 受賞対象分野:情報科学)を受賞しました。
riken.jp
🏆 I'm delighted to share that I've won a 2025 COPSS Emerging Leader Award! 😃 And congratulations to my fellow winners! 🙌🏽 Check out how each of us is improving and advancing the profession of #statistics and #datascience here: tinyurl.com/copss-emerging-leader-award
Congratulations to the 2025 #COPSS Awardees, @ericjdaza.com, @lucystats.bsky.social, @lestermackey.bsky.social, and all of you. I hope to congratulate you at #JSM2025 in Nashville with @amstatnews.bsky.social. God I hope to go. #rstats #statssky
🏆 I'm delighted to share that I've won a 2025 COPSS Emerging Leader Award! 😃 And congratulations to my fellow winners! 🙌🏽 Check out how each of us is improving and advancing the profession of #statistics and #datascience here: tinyurl.com/copss-emerging-leader-award
NeurIPS 2025 is soliciting self-nominations for reviewers and ACs. Please read our blog post for details on eligibility criteria, and process to self-nominate:
Self-nomination for reviewing at NeurIPS 2025 – NeurIPS Blog
Communications Chairs 2025 2025 Conference
blog.neurips.cc
Congratulations to @lestermackey.bsky.social for receiving the 2025 COPSS Award! 🎉👏 Lester is currently the Chair of the Section on Bayesian Statistical Sciences (SBSS) of the American Statistical Association.
Off to #AAAI25! We're presenting #SatCLIP (w/ @marccoru.bsky.social, @estherrolf.bsky.social, @calebrob6.bsky.social & @lestermackey.bsky.social) at the 12.30-2.30pm poster session on Feb 28! Let me know if you're around & want to chat #GeoAI!🛰️ Paper: tinyurl.com/5eejz5kw Code: tinyurl.com/2zm64967
I've been shocked that a theory-driven method yields practical results this good, especially on attention approximation. I proposed my best new optimizer design originally as a dumb baseline; the fact that you can get these efficiency gains with a principled approach makes me a lil insecure.
New guarantees for approximating attention, accelerating SGD, and testing sample quality in near-linear time