BenjMurrell

@benjmurrell.bsky.social

🇸🇪🇿🇦 Researcher at Karolinska. Comp bio. Phylogenetics. Deep learning. All in Julia. Virology. Immunology. https://scholar.google.com/citations?user=I80vy5cAAAAJ

A technical thread on loss scaling in diffusion and flow matching models (related to a new preprint): Since the dawn of time, people have been messing with (or dropping entirely) these pesky time-dependent loss scaling terms, mostly because the models train better without them.

Bild

We figured out flow matching over states that change dimension. With "Branching Flows", the model decides how big things must be! This works wherever flow matching works, with discrete, continuous, and manifold states. We think this will unlock some genuinely new capabilities.

My lab, at Karolinska, in Stockholm, is looking for a PhD student with a computational/quantitative background to work on probabilistic/generative models of proteins (structure and sequence). The research will involve methods development, and applications in vaccine design.

Transformer/attention folks: I saw massive activations in Qwen's keys, always towards the end of each head, especially in layer 1. Turns out this is directly driven by the key projection bias (which Qwen has but eg. Llama3 does not). These large values are where RoPE has the slowest(?) effect. Why?

Attention key bias, layer 1.