Despite the huge amount of training data, ESM-C is still unable to distinguish mean-pooled CDRH3 representations of real antibodies from those of antibodies w/ scrambled CDRs. Only VHHBERT seems to do this, and only for natural sequences
Diego del Alamo
@delalamo.xyz
Computational protein engineering & synthetic biochemistry at Takeda Opinions my own https://linktr.ee/ddelalamo
LinkedIn as an employer looking for candidates is like ipTM: great for filtering candidates discovered elsewhere, but absolute dogshit for discovery. Quite a few people find ways to fluff up mostly-empty CVs and post a firehose of AI-generated posts. Anyway, role still open
Takeda Pharma is hiring scientists at both the straight-out-of-PhD and senior levels in the AIML group, particularly those with experience training, modifying, and applying foundation models. Please reach out if youre interested and want to make the jump to industry, link below👇
Indeed.. here is "generate the first page of a great Nature paper that will get lots of citations. It must be very novel."
Takeda Pharma is hiring scientists at both the straight-out-of-PhD and senior levels in the AIML group, particularly those with experience training, modifying, and applying foundation models. Please reach out if youre interested and want to make the jump to industry, link below👇
Does anyone know what paper this figure is from? The citation I have for it turned out to be wrong
PDB 7E5Y is a funny case of two copies of the same protein in the same asymmetric unit, one of which has a register shift error (chain C, green) while the other does not (chain H, orange). Here's conserved J-gene Trp 118 (IMGT; 108 in PDB). Trastuzumab shown in teal for reference
The proteina-complexa paper is the 2nd this month suggesting that high K+E content, a staple of ProteinMPNN designs, is predictive of poor expression (the bits in bio benchmark paper showed a few weeks ago). In their case, they measure poor sequence recovery in phage display
This is new? Instead of relying on intermediate estimates of the denoised state, Proteina-Complexa just does the full denoising, calculates whatever it needs to, and sends the info back to continue guiding the unfinished diffusion roll-out
🔸 Quantitatively, our inference-time scaling strategies (we use, for instance, MCTS, Feynman-Kac Steering and Beam Search) outperform previous hallucination methods under normalized compute budgets, setting a new state-of-the-art in in-silico binder design. (9/n)
Sometimes you hit a home run, and sometimes you get to first base
Statistically significant chuckles: who is using humour at scientific #conferences? #ProcB #BiologicalSciencePractices #TheoreticalBiology royalsocietypublishing.org/rspb/article...
So I can't say I've ever seen residual cross-attention before (where the final representations attend to earlier representations of the input data); is there any literature on when and where to use this?
Alex Morehead, et al.: Zatom-1: A Multimodal Flow Foundation Model for 3D Molecules and Materials https://arxiv.org/abs/2602.22251 https://arxiv.org/pdf/2602.22251 https://arxiv.org/html/2602.22251
Diffusion models actually learn a series of time-indexed energy landscapes, which are corrupted with different amounts of noise. The ranking ability of ProteinEBM peaked slightly above t=0. Inspired by this finding we trained an "expert" model only on low time levels, which we call ProteinEBM-x.
You can triple your hit rate without changing your NN by simply using a different search algorithm during sampling
We've been investing heavily in better protein language models (PLMs), but relatively little work addresses how to best generate with them. We present a new search-based method for PLMs and exhaustively benchmark models and methods, including with in vitro data from antibody therapeutics campaigns.🧵
> paper claims astonishing progress on protein folding problem > ask if it’s interesting proteins or villin headpiece denatured with urea > they say results are robust to protein sequence properties > open the pdf > villin headpiece in urea
I wrote a blog post about the future of structural bioinformatics. Where to go after AlphaFold? How do we avoid the field becoming a load of half-baked LLMs? Let me know what you think. jgreener64.github.io/posts/struct...
Where next for structural bioinformatics?
jgreener64.github.io
One of the most interesting parts of this workflow is the "sunk cost fallacy" estimator that predicts how promising a particular mutational line of inquiry is, and whether it is worth abandoning in favor of others
What comes after de novo? Automated lead optimization of proteins with CRADLE-1 [new] Automated multi-property lead optimization of diverse protein modalities by fine-tuning protein language models with lab-in-the-loop data.
Most benchmarks for drug discovery AI don't effectively evaluate generative models; instead, due to the data's incompleteness, they rely on surrogate functions, like fwd folding, property prediction, or ranking previously characterized designs. No idea what the solution is here
New paper from former PhD student @tkschulze.bsky.social on supervised learning of protein variant effects across large-scale mutagenesis datasets MAVE/DMS experiments provide large amounts of data for benchmarking variant effect predictors, but may be difficult to use in supervised learning. 1/5
I regret to inform you that if you are job hunting on LI and this is your profile pic then you’re ngmi
RIP my LinkedIn after mentioning a job opening in my group ☠️
The two co-first authors of this research paper, Yang & Yang, has decided to sort their names alphabetically
Protenix trained an identical model with way more training data (2025 cutoff instead of 2021), demonstrating that antibody-antigen modeling, but not protein-ligand modeling, is currently data-limited (DQ SR % means % DockQ≥0.23) with this architecture
Scribbled on at least one whiteboard in every office with "do not erase" written below
SGkhIFlvdSBjcmFja2VkIG15IGFkbWl0dGVkbHkgbm90LXZlcnktY29tcGxleCBjb2RlLCB3ZWxsIGRvbmUhIEknbSBsb29raW5nIGZvciBhIG5ldyByb2xlIGRvaW5nIGRhdGEgc3Rvcnl0ZWxsaW5nIHdvcmssIGlkZWFsbHkgZm9yIG5vbi1VSyBhbmQgbm9uLVVTIG5ld3NwYXBlcnMuIE1heWJlIHlvdSBrbm93IG9mIHNvbWV0aGluZy4uLj8gUGxlYXNlIERNIG1lIQ==
I have a concern with this paper and I want someone who knows more than me to confirm if it is founded or not. The title makes a pretty specific claim about epistasis predictions, but the method does not seem sound for masked LMs (1/3)
Beyond additivity: zero-shot methods cannot predict impact of epistasis on protein properties and function [new] models capture single/non-epistatic effects but miss complex epistatic interactions on protein properties and function.
Hot take: this "chain X[auth Y]" notation on FASTA files pulled from the PDB is needlessly confusing, adds nothing, and needed to be changed yesterday
This plot is quite the indictment of fine-tuned PLMs, showing how performance is entirely data-dependent and, at the upper end of performance, equally achievable with randomized model weights
We made FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns.
We're hiring - looking for folks with experience designing and training foundation models from scratch, particularly biomolecular language models, GNNs, or vision models. Lots of room for creativity in this role. Contact me if you have any interest
Research Senior Scientist AI/ML Foundational Models at Takeda Pharmaceutical
Learn more about applying for Research Senior Scientist AI/ML Foundational Models at Takeda Pharmaceutical
jobs.takeda.com