Tony S.F.

@tonysf.bsky.social

Ass. Prof. of AI at CentraleSupélec in the Centre pour la Vision Numérique.

Big lack of qualified reviewers? The cost of generation seems headed to 0 relative to the cost of verification; seems obvious that we must change norms so verification is seen as more of a contribution? What if we require authors to review for some number of conferences before being able to submit?

photo taken from https://x.com/mar_kar_/status/2074240160758444372

Bienvenue à Nice, ville avec une des plus basses températures de France ! (ps: je recrute potentiellement un.e postdoc sur 24 mois d'ici fin 2026 sur des questions, plutôt théoriques, liées aux LLMs en post-training / alignment, je dis ça, je dis rien)

Bild

i got asked by a friend if my figures were made with chatgpt because he liked them and, while for this time i could say no and show him a different talk with the same figures from before chatgpt, it saddened me to think everyone will likely assume this is the case from now on

Dirk Lorenz @dirque.bsky.social · 2mo ago

I noticed that LLMs already affect student presentations a lot, but now at #SIAMOP26 I realize that also here, people use ChatGPT et al. to produce their slides (at least, that's the obvious explanation for all the "Why this matters" boxes I see on slides...)

New paper! We analyze proximal preconditioned gradient methods that extend Muon/Scion to handle nonconvex constraints (Stiefel manifold, spectral sphere, norm balls, ...) with convergence guarantees under heavy-tailed noise + variance reduction w/ STORM! arxiv.org/abs/2605.11850

BildBildBild

I heard that it's easier to get an h100 on Jean Zay than an a100, kind of funny. The hour multiplier for consumption (i.e. one h100 hour costs 4 credits) should take into account demand.

you can improve your collaborators' writing clarity by being too dumb to fill in the gaps of what they've written, and arguing it must be wrong until they write it clearly enough that even you can understand.

In conditional gradient sliding you are using the conditional gradient algorithm to "chase" the projected Nesterov algorithm. Instead of computing the projection, you do some conditional gradient steps to approximate it. I wonder if you can do the same with FISTA/accelerated proximal point alg ?

nerd sniped by the bayesian learning rule again and still unsatisfied... ok, so you can explain a lot of DL optimization algorithms with certain approximations of various posteriors but that's kind of kicking the can down the road - the question becomes: why those approximations instead of others?

My paper on Generalized Gradient Norm Clipping & Non-Euclidean (L0, L1)-Smoothness (together with collaborators from EPFL) was accepted as an oral at NeurIPS! We extend the theory for our Scion algorithm to include gradient clipping. Read about it here arxiv.org/abs/2506.01913

Bild