Nicola Branchini

@nicolabranchini.bsky.social

🇮🇹 ProbAI Research Fellow @warwickstats.bsky.social. Previously @ellis.eu Stats PhD @edinunimaths.bsky.social @aalto.fi. 🤔💭 about Monte Carlo, approximate inference, UQ

"We're all worried," as what it means to do research (in my field, Theoretical CS) seems to be shifting, and shifting fast. What to do? Senior researchers must lead by example, knowing that not everything will pan out. What I'm suggesting below may not work everywhere, but here's my own advice: 1/

So much masochism in academia. We try to find reasons to reject (I do try not to). We decide to reject a paper for reason X & Y when multiple, very closely related work has been published at the same venue (or better for some, whatever that means), where X&Y could have equally applied as criticism

Federic Perlino, who is currently in the second year of a PhD working with Theo Damoulas and I has just arxived the first paper arising from it: arxiv.org/abs/2607.09645. In which he develops models for function composition over graphical structures using Gaussian processes.

Deep Gaussian Processes on Directed Acyclic Graphs

Many real-world processes can be represented as compositions of functions along a directed acyclic graph (DAG). In causal modelling, these correspond to the underlying mechanisms; in engineering, to m...

arxiv.org

Things I am not interested in, re: your research talk - how prestigious is the venue where you / the related work published - how much stuff you know / you did Things I am interested in: - the core ideas (explained with intent to teach, not impress, ideally) - how it fits within broader landscape

🧵 Preprint alert! Introducing Augmented Gaussian sum filters (AGSF), a novel class of Bayesian filtering algorithms which unifies Gaussian sum (GSF) and particle filters (PF) by interpolating continuously between them, while being robust to common failure modes. arxiv.org/abs/2605.21698

A Gaussian Sum Filter for Unifying Gaussian and Particle Filters

State-space models (SSMs) are a broad class of probabilistic models for dynamical systems with many applications in engineering and science. Bayesian filtering is analytically tractable only in the li...

arxiv.org

"Accept (spotlight)" at ICML'26 😎 Our paper brings particle filters back to life: autoregressive diffusion models + posterior sampling yield optimal proposals for Bayesian filtering, scaling up to GenCast-sized systems. arxiv.org/abs/2605.20028 w/ Thomas Savary and @francois-rozet.bsky.social

Training-Free Bayesian Filtering with Generative Emulators

Bayesian filtering is a well-known problem that aims to estimate plausible states of a dynamical system from observations. Among existing approaches to solve this problem, particle filters are theoret...

arxiv.org

ICML Conference@icmlconf.bsky.social · 3mo ago

Congrats again to authors of accepted #ICML2026 papers! The camera-ready deadline is 5/28. Drawing your attention to two specific features: 1. As last year, to help communicate research to a broad audience, papers will have lay summaries. Tips & details in blog 1/3

4 ICLR papers 🥳 There’s an insightful story between them: If you sample LLMs multiple times, they are calibrated, even on higher levels [1], but they cannot talk about this uncertainty in a single prompt [2], so you have to help them out to gather information Bayes-optimally [3]

Bild

I want to advertise the PhD thesis of my good friend and luckily also collaborator Thomas Guilmeau theses.hal.science/tel-05474635/ I've learnt so much talking to Thomas about divergence minimization.

Divergence-minimization for variational inference, black-box global optimization, and importance sampling

Many methods across applied mathematics aim at constructing specific parametric probability distributions. Examples of these tasks include evolution strategies or simulated annealing for black-box global optimization, Monte Carlo methods based on adaptive importance sampling, and variational inference algorithms in machine learning. The construction of such distributions can often be formulated as the minimization of a statistical divergence. However, these divergence-minimization problems are challenging because of the following reasons. First, the specific geometry of the considered set of parametric probability distributions needs to be taken into account. Second, efficient evaluations of statistical divergences come with important noise. Third, divergence-minimization problems are generally non-convex. Because of these difficulties, standard methods may fail to converge to good solutions. We tackle these challenges in this thesis.First, we show that evolution strategies, which are sampling-based algorithms for black-box global optimization problems, can be analysed through the lens of divergence-minimization problems. Our approach allows to establish and quantify the improvement brought at each iteration of the algorithms. We show that existing methods fit within our framework, yielding a new approach for their analysis. We also establish improvement results for two novel algorithms, one related with mixture models, and another one using heavy-tailed parametric probability distributions.Second, we consider the minimization of a regularized Rényi divergence over an exponential family. We propose to solve this problem with a stochastic Bregman proximal-gradient algorithm, with biased gradient estimator. By leveraging the geometry of the exponential family, we prove strong convergence guarantees for our algorithm, with proof techniques that are of interest beyond the considered problem. We then extend this algorithm to propose an adaptive simulated annealing algorithm with solid theoretical understanding. We show through a rigorous benchmarking that our algorithm outperforms similar non-adaptive algorithms.Finally, we go beyond exponential families and look at variational inference problems over lambda-exponential families. Using generalized convexity tools, we give new sufficient optimality conditions for these problems, which generalize existing similar results for the exponential family. For the resolution of these problems, we propose novel proximal-like algorithms that exploit the geometry underlying the lambda-exponential family. These results are especially useful for heavy-tailed distributions. We then leverage our results to propose an adaptive importance sampling algorithm to cover these cases. We show that our algorithm is able to learn Student distributions that capture the location, scale, and tail behaviour of target distributions, both in heavy-tailed and light-tailed cases.

theses.hal.science

I disagree with the view that peer review isn't problematic just because your papers usually get accepted. Big difference between being accepted and being accepted for the right reasons (nevermind having proper feedback)

I hate when people refrain from giving me blunt feedback on my work out of politeness. I really want to know if you don't see the point 😄 I promise you can't hurt my feelings. This doesn't happen often, but more so at conferences than anywhere else.