Han Bao

@han-b.bsky.social

Associate Professor@The Institute of Statistical Mathematics, working in machine learning theory https://hermite.jp/

Last week I introduced this work in a domestic workshop among mathematicians and physicists, and one astrophysicist gave me a new insight to interpret the self-stabilizing potential! This is a real virtue of having a workshop with people from the communities next to us.

Han Bao@han-b.bsky.social · 2w ago

I spent the whole weekend to understand the crux of self-stabilization (arxiv.org/abs/2209.15594). To me, the most interesting part of this paper is how they model the edge of stability by ODE---because the stepsize is no longer infinitesimal at EoS and it is apparently not trivial to derive an ODE

それまで人件費として支払われていたコストがAIの労働代替によって海外の寡占的なテック企業にサービスフィーとして吸い上げられ、各国の企業が雇用を通じて社会に提供していた「消費者と需要を生み出す機能」が毀損されてゆくという問題について、もっと真剣に議論されるべきだと思います。 そっちのほうがLLMはAIかどうかなんて神学論争よりずっと重要。

スドー@stdaux.bsky.social · 2w ago

「人を雇うよりAIのほうが安い」という場面も多いだろうが、問題は人を雇うと従業員の給与となって誰かの食費や学費に化けて国内消費に回るのに対して、AI使用料はどこかに消えるというのがな

I spent the whole weekend to understand the crux of self-stabilization (arxiv.org/abs/2209.15594). To me, the most interesting part of this paper is how they model the edge of stability by ODE---because the stepsize is no longer infinitesimal at EoS and it is apparently not trivial to derive an ODE

Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability

Traditional analyses of gradient descent show that when the largest eigenvalue of the Hessian, also known as the sharpness $S(θ)$, is bounded by $2/η$, training is "stable" and the training loss decre...

arxiv.org

Can't agree more. Previously one guy told me that if you want your theory paper to get accepted, the best thing is to write 30+ pages appendix (to overwhelm!). That's insane. Unless truly "necessary" (I know it's debatable), we should compress/distill formal proofs. Then we get better intuition.

People used to be able to impress and intimidate reviewers with complicated proofs. This will change. In the age of AI inscrutable proofs are cheap. It is understandable proofs that are valuable. Opaque complexity is now it is a sign of laziness or lack of insight.

GD is known to have max-margin bias, but its convergence rate is extremely slow. However, we can observe fairly decent convergence behaviors in practice as seen in the figure. I have been thinking this for a while and figured out early-stage weak convergence is possible! arxiv.org/abs/2608.04382

Bild

And the new initial meta review system is also stressful than I expected. I supposed it'd reduce rebuttal workload for both author and reviewer sides, but actually just gave us an extra burden. Very few ppl care the initial meta review and ppl do endlessly long rebuttals until the reviewers gave up🤦‍♂️

Really stressful to see AI-driven discussions during the review period this year... as an AC, I really don't have an idea how to facilitate discussion among AI-ish paper's authors and AI-ish reviewers (and the worst thing is that we cannot suppose they are truly AI even if it's very likely)

AI is an interesting mechanism: it leads some scientists to reveal what they would not admit willingly, that is, that they actually don't care at all about scientific rigour. AI is not the problem here, it just makes the problem visible...

While LLM/diffusion are extremely popular, we still can find classical topics in the poster session. This paper proposes convexified heterogeneous OT (unlike nonconvex Gromow-Wasserstein). In essence, | E[d(x,X)|y] - E[d(y,Y)|x] | is regarded as the base cost for OT. arxiv.org/abs/2606.02047

Convex Distance Operator Transport: A Convex and Geometry-Preserving Formulation

We introduce Convex Distance Operator Transport (CDOT), the first convex optimal transport framework that aligns distributions across heterogeneous domains by jointly preserving feature correspondence...

arxiv.org

On Monday @ncollina.bsky.social @iraglobusharris.bsky.social and I are giving a tutorial at ICML on (multi)calibration and its applications. You can find slides and an annotated bibliography on the website: calibration-tutorial.github.io as well as an interactive demo of online calibration algs.

Calibration, Decisions, and Collaboration in Learning | ICML 2026

An ICML 2026 tutorial on making probabilistic predictions trustworthy for downstream decision-making and collaboration.

calibration-tutorial.github.io

My first PhD student Xianliang worked hard out this: In Muon, polar decomposition should always precede momentum, which significantly improves signal recovery. I'm excited to share this since few theory has been working on the benefit of momentum in Muon! arxiv.org/abs/2606.03899

Denoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral Filtering

Muon has recently demonstrated strong empirical performance in large language model training, but the theoretical role of momentum in Muon remains unclear. Existing analyses of Muon either remove mome...

arxiv.org

#AISTATS2026 It was extremely great to first see someone whose paper I have closely read, those with whom I collaborated recently without having met in-person, got invited to a next workshop, etc. Even though the time is quite challenging before the deadline😅, I really enjoyed the conference!

My new policy: if someone asks me to read something, I ask them how they used AI in creating it, and what validation/processing they applied to the AI outputs. (I disclose the same.) Been burned by giving too much attention to (undisclosed) slop folks have sent me...

D'ailleurs, j'ai pu profiter de mon séjour à Paris cette fois-ci pendant l'escale avant d'aller au Maroc. J'ai vu quelques endroits que j'aime là-bas---surtout le Centre de Pompidou, même s'il est en rénovation---aprés 7 ans! Maintenant, c'est le moment de se concentrer intensément sur le travail...

Bild

Tomorrow is the very first class of my lecture at ISM (I'm gonna introduce learning theory and convex analysis). It's extremely useful for myself as well to review bunch of facts and proofs, but I need to rush because I've prepared for only half a semester😅

Bild

Sad to see more and more people rely on LLMs to generate reviews... (for some reasons, it is very easy to find at a glance; and senior people tend to rely, tho I don't intend to generalize this)