Diego del Alamo

@delalamo.xyz

Computational protein engineering & synthetic biochemistry at Takeda Opinions my own https://linktr.ee/ddelalamo

Despite the huge amount of training data, ESM-C is still unable to distinguish mean-pooled CDRH3 representations of real antibodies from those of antibodies w/ scrambled CDRs. Only VHHBERT seems to do this, and only for natural sequences

BildBildBild

LinkedIn as an employer looking for candidates is like ipTM: great for filtering candidates discovered elsewhere, but absolute dogshit for discovery. Quite a few people find ways to fluff up mostly-empty CVs and post a firehose of AI-generated posts. Anyway, role still open

Diego del Alamo@delalamo.xyz · 3mo ago

Takeda Pharma is hiring scientists at both the straight-out-of-PhD and senior levels in the AIML group, particularly those with experience training, modifying, and applying foundation models. Please reach out if youre interested and want to make the jump to industry, link below👇

Takeda Pharma is hiring scientists at both the straight-out-of-PhD and senior levels in the AIML group, particularly those with experience training, modifying, and applying foundation models. Please reach out if youre interested and want to make the jump to industry, link below👇

PDB 7E5Y is a funny case of two copies of the same protein in the same asymmetric unit, one of which has a register shift error (chain C, green) while the other does not (chain H, orange). Here's conserved J-gene Trp 118 (IMGT; 108 in PDB). Trastuzumab shown in teal for reference

Bild

The proteina-complexa paper is the 2nd this month suggesting that high K+E content, a staple of ProteinMPNN designs, is predictive of poor expression (the bits in bio benchmark paper showed a few weeks ago). In their case, they measure poor sequence recovery in phage display

Bild

This is new? Instead of relying on intermediate estimates of the denoised state, Proteina-Complexa just does the full denoising, calculates whatever it needs to, and sends the info back to continue guiding the unfinished diffusion roll-out

Bild
Karsten Kreis@karstenkreis.bsky.social · 5mo ago

🔸 Quantitatively, our inference-time scaling strategies (we use, for instance, MCTS, Feynman-Kac Steering and Beam Search) outperform previous hallucination methods under normalized compute budgets, setting a new state-of-the-art in in-silico binder design. (9/n)

So I can't say I've ever seen residual cross-attention before (where the final representations attend to earlier representations of the input data); is there any literature on when and where to use this?

Bild
arXiv cs.LG Machine Learning@cslg-bot.bsky.social · 5mo ago

Alex Morehead, et al.: Zatom-1: A Multimodal Flow Foundation Model for 3D Molecules and Materials https://arxiv.org/abs/2602.22251 https://arxiv.org/pdf/2602.22251 https://arxiv.org/html/2602.22251

Diffusion models actually learn a series of time-indexed energy landscapes, which are corrupted with different amounts of noise. The ranking ability of ProteinEBM peaked slightly above t=0. Inspired by this finding we trained an "expert" model only on low time levels, which we call ProteinEBM-x.

Bild

> paper claims astonishing progress on protein folding problem > ask if it’s interesting proteins or villin headpiece denatured with urea > they say results are robust to protein sequence properties > open the pdf > villin headpiece in urea

One of the most interesting parts of this workflow is the "sunk cost fallacy" estimator that predicts how promising a particular mutational line of inquiry is, and whether it is worth abandoning in favor of others

Bild
AI x Bio Discovery@aixbiobot.bsky.social · 5mo ago

What comes after de novo? Automated lead optimization of proteins with CRADLE-1 [new] Automated multi-property lead optimization of diverse protein modalities by fine-tuning protein language models with lab-in-the-loop data.

What comes after de novo? Automated lead optimization of proteins with CRADLE-1

Most benchmarks for drug discovery AI don't effectively evaluate generative models; instead, due to the data's incompleteness, they rely on surrogate functions, like fwd folding, property prediction, or ranking previously characterized designs. No idea what the solution is here

Protenix trained an identical model with way more training data (2025 cutoff instead of 2021), demonstrating that antibody-antigen modeling, but not protein-ligand modeling, is currently data-limited (DQ SR % means % DockQ≥0.23) with this architecture

Bild

I have a concern with this paper and I want someone who knows more than me to confirm if it is founded or not. The title makes a pretty specific claim about epistasis predictions, but the method does not seem sound for masked LMs (1/3)

AI x Bio Discovery@aixbiobot.bsky.social · 6mo ago

Beyond additivity: zero-shot methods cannot predict impact of epistasis on protein properties and function [new] models capture single/non-epistatic effects but miss complex epistatic interactions on protein properties and function.

Beyond additivity: zero-shot methods cannot predict impact of epistasis on protein properties and function