Alex Thiery

@alexxthiery.bsky.social

Associate Prof. in ML & Statistics at NUS 🇸🇬 MonteCarlo methods, probabilistic models, Inverse Problems, Optimization https://alexxthiery.github.io/

Here's how the gradient flow for minimizing KL(pi, target) looks under the Fisher-Rao metric. I thought some probability mass would be disappearing on the left and appearing on the right (i.e. teleportation), like a geodesic under the same metric, but I was very wrong... What's the right intuition?

<proud advisor> Hot off the arXiv! 🦬 "Appa: Bending Weather Dynamics with Latent Diffusion Models for Global Data Assimilation" 🌍 Appa is our novel 1.5B-parameter probabilistic weather model that unifies reanalysis, filtering, and forecasting in a single framework. A thread 🧵

Cute way to upper bound the connective constant of Z^d. For some length L, enumerate {w_1, w_2, ... , w_N} the Self-Avoiding-Walks of size L. An upper bound is given by the largest eigenvalue of the NxN matrix where M_{i,j}=1 iff there is a SAW of size (L+1) that starts with w_i and ends with w_j.

Bild
Alex Thiery@alexxthiery.bsky.social · last yr.

Approximating N(L), the number of Self-Avoiding-Walks in Z^2 of length L, is an assignment in my Simulation course this year. The connective constant is: C = \lim N(L)^1/L ~ 2.638.. Still open-problem to this day: is it true that 1/C equals the zero of the polynomial P(x)=581*x^4 + 7*x^2 - 13 😱

Approximating N(L), the number of Self-Avoiding-Walks in Z^2 of length L, is an assignment in my Simulation course this year. The connective constant is: C = \lim N(L)^1/L ~ 2.638.. Still open-problem to this day: is it true that 1/C equals the zero of the polynomial P(x)=581*x^4 + 7*x^2 - 13 😱

Bild

Tonight, on the taxi ride home, the 72-year-old driver, super friendly and insightful, spent ~10 minutes sharing his first impressions when using DeepSeek, comparing its pros and cons with ChatGPT, and even diving into the potential geopolitical implications 😅

When implementing parallel tempering, it's fashionable to alternate even and odd index temperature swap to try to maximise the inter-temperature movements. But when the temperatures are appropriately tuned, this very new paper by Roberts & Rosenthal shows that the gains are quite modest!

Bild

Asked to the students of my "statistical simulation" class: In Buffon's experiment where a needle of length L falls on parallel strips (unit width), the needle crosses 2L/π strips on average. To maximize the accuracy of the resulting estimate of π, how should one choose the length of the needle?

A perfectly trained diffusion/flow model will just memorize the training data, so why isn't it the case in practice? Super interesting work 👇👇

Surya Ganguli@suryaganguli.bsky.social · 2y ago

Our new paper! "Analytic theory of creativity in convolutional diffusion models" lead expertly by @masonkamb.bsky.social arxiv.org/abs/2412.20292 Our closed-form theory needs no training, is mechanistically interpretable & accurately predicts diffusion model outputs with high median r^2~0.9

I've been thinking about in-context learning for nearly 3 years. While there is still plenty I don't fully understand, five papers have--to a very large extent--shaped my perspective on it, and I believe everyone should read them.