Tony S.F.
@tonysf.bsky.social
Ass. Prof. of AI at CentraleSupélec in the Centre pour la Vision Numérique.
It's surely just regurgitating from the many counterexamples to the Jacobian Conjecture that were in the training data
Well, the Jacobean Conjecture appears to have just been proven false by Fable.
Slides for our ICML tutorial on Memorization and Generalization of Diffusion and Flow Matching Models are now available ! 🌀 memorization-generalization.github.io @quentinbertrand.bsky.social
ICML 2026 Tutorial - Generalization and Memorization in Flow Matching and Diffusion
memorization-generalization.github.io
Big lack of qualified reviewers? The cost of generation seems headed to 0 relative to the cost of verification; seems obvious that we must change norms so verification is seen as more of a contribution? What if we require authors to review for some number of conferences before being able to submit?
Bienvenue à Nice, ville avec une des plus basses températures de France ! (ps: je recrute potentiellement un.e postdoc sur 24 mois d'ici fin 2026 sur des questions, plutôt théoriques, liées aux LLMs en post-training / alignment, je dis ça, je dis rien)
i got asked by a friend if my figures were made with chatgpt because he liked them and, while for this time i could say no and show him a different talk with the same figures from before chatgpt, it saddened me to think everyone will likely assume this is the case from now on
I noticed that LLMs already affect student presentations a lot, but now at #SIAMOP26 I realize that also here, people use ChatGPT et al. to produce their slides (at least, that's the obvious explanation for all the "Why this matters" boxes I see on slides...)
my coauthors have convinced me that it's not the best decision to name our NonSmooth Frank-Wolfe algorithm NSFW... i thought it was catchy.
New paper! We analyze proximal preconditioned gradient methods that extend Muon/Scion to handle nonconvex constraints (Stiefel manifold, spectral sphere, norm balls, ...) with convergence guarantees under heavy-tailed noise + variance reduction w/ STORM! arxiv.org/abs/2605.11850
Also, a shoutout to this amazing paper by @tonysf.bsky.social and collaborators, which is well worth reading: arxiv.org/abs/2502.07529
Training Deep Learning Models with Norm-Constrained LMOs
In this work, we study optimization methods that leverage the linear minimization oracle (LMO) over a norm-ball. We propose a new stochastic family of algorithms that uses the LMO to adapt to the geom...
arxiv.org
A new paper about how to scale your training of LLMs when increasing the token budget, based on the convergence theory! Lots of empirical experiments validating the assumptions we make. arxiv.org/abs/2603.21191
On the Role of Batch Size in Stochastic Conditional Gradient Methods
We study the role of batch size in stochastic conditional gradient methods under a $μ$-Kurdyka-Łojasiewicz ($μ$-KL) condition. Focusing on momentum-based stochastic conditional gradient algorithms (e....
arxiv.org
So when you're doing muon with weight decay to train nanoGPT you're using frank-wolfe to train a frank-wolfe machine
I missed this post but it is pure gold. www.colincornaby.me/2025/08/in-t...
In the Future All Food Will Be Cooked in a Microwave, and if You Can’t Deal With That Then You Need to Get Out of the Kitchen
Update 8/8/2025 – I wrote this the day before a certain post by a popular developer services company. I’ve seen some comments this is a rebuttal – it wasn’t meant to be! But…
colincornaby.me
I heard that it's easier to get an h100 on Jean Zay than an a100, kind of funny. The hour multiplier for consumption (i.e. one h100 hour costs 4 credits) should take into account demand.
Come check out our #ICCV2025 poster for "Multi-modal Identity Extraction" at (Exhibit Hall I #73).
you can improve your collaborators' writing clarity by being too dumb to fill in the gaps of what they've written, and arguing it must be wrong until they write it clearly enough that even you can understand.
I started to read this paper arxiv.org/abs/2510.17503 and I thought huh the analysis is so much like Frank-Wolfe, then I remembered that Frank-Wolfe and DC algorithms are dual. Probably, a Frank-Wolfe god like Jaggi knows that but it's not mentioned in the paper; I must be missing something simple.
arxiv.org
Have you ever written a paper, and you see a small variation you could easily cover with your analysis etc but you don't do it? But you know if someone else did it right after, you would be upset you didn't include it? It happened to me again today! arxiv.org/abs/2510.16468
arxiv.org
Abbas Khademi, Antonio Silveti-Falls Adaptive Conditional Gradient Descent https://arxiv.org/abs/2510.11440
Now accepted at #NeurIPS2025 :)
📣 New preprint 📣 **Differentiable Generalized Sliced Wasserstein Plans** w/ L. Chapel @rtavenar.bsky.social We propose a Generalized Sliced Wasserstein method that provides an approximated transport plan and which admits a differentiable approximation. arxiv.org/abs/2505.22049 1/5
In conditional gradient sliding you are using the conditional gradient algorithm to "chase" the projected Nesterov algorithm. Instead of computing the projection, you do some conditional gradient steps to approximate it. I wonder if you can do the same with FISTA/accelerated proximal point alg ?
nerd sniped by the bayesian learning rule again and still unsatisfied... ok, so you can explain a lot of DL optimization algorithms with certain approximations of various posteriors but that's kind of kicking the can down the road - the question becomes: why those approximations instead of others?
My paper on Generalized Gradient Norm Clipping & Non-Euclidean (L0, L1)-Smoothness (together with collaborators from EPFL) was accepted as an oral at NeurIPS! We extend the theory for our Scion algorithm to include gradient clipping. Read about it here arxiv.org/abs/2506.01913
Found this on r/math: priority dispute in pure math that has come to a head, arxiv.org/abs/2507.20816
History of the canonical basis and crystal basis
The history of the canonical basis and crystal basis of a quantized enveloping algebra and its representations is presented
arxiv.org