AlphaFold database has entered the era of complexes. Together with NVIDIA, DeepMind and EBI, we use ColabFold, OpenFold and MMseqs2-GPU to predict ~31 million complexes (homo & hetro-dimers) resulting in 1.8 million high-quality predictions 📄 research.nvidia.com/labs/dbr/ass... 🌐 alphafold.ebi.ac.uk
Christian Dallago
@machine.learning.bio
🏳️🌈 NVIDIA & Duke. Was Allianz, VantAI, TUM. BioCS+ML dude. Lab page: https://machine.learning.bio GScholar: https://scholar.google.com/citations?user=4q0fNGAAAAAJ
You asked, we listened. Millions of AI-predicted protein complex structures are now available in the #AlphaFold Database. This spans homodimers from 20 of the most studied species, including humans, as well as the World Health Organization’s priority pathogens list. www.ebi.ac.uk/about/news/t...
Five years ago, we released FLIP. The core question was: can ML models for protein fitness prediction generalize in the ways that actually matter for protein engineering, i.e. low data, extrapolation to more mutations, out-of-distribution sequences?
We made FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns.
Our latest protein family-based GenAI collection of tools and datasets, ProFam, is out now. Everything -- from data, training and inference code, to a 215M llama-based ProFam-1 are fully open sourced. 🧵
Built by CATH, TÜM and NVIDIA, ProFam-1 is our new open-source protein family language model (pfLM) designed to generate functional protein variants and predict fitness using in-context example sequences.
Another exciting opportunity, this time as a colleague at Duke! Join as tenure track assistant prof. in Cell Bio & let’s work on closing the gap between in-silico and in-vivo: www.nature.com/naturecareer... Important: application closes Nov 1st!!!
Tenure-Track Assistant Professor Position –AI/ML for Cell Biology - Durham, North Carolina (US) job with Duke University School of Medicine | 12844591
Tenure-Track Assistant Professor Position –AI/ML for Cell Biology
nature.com
Another opening: Senior Multiscale Biology Applied Research Scientist! nvidia.eightfold.ai/careers/job/... Are fascinated by fundamental data modalities across biology like RNA-seq, mass spec & want to build computational tools that harnessing data to build intelligence? Come: join the team!
Senior Applied Research Scientist, Multiscale Biology | NVIDIA Corporation
Apply your expertise in engineering biology through algorithms and tools for genes, tissues, organisms, and populations. Conduct collaborative applied research in multiscale biology using deep learnin...
nvidia.eightfold.ai
Are you passionate about leading collaborative, fast moving, applied bioinformatics research projects that help the entire community move forward? Apply to work in my team at NVIDIA: nvidia.eightfold.ai/careers/job/...
Senior Applied Research Scientist, Bioinformatics | NVIDIA Corporation
Lead applied and collaborative research programs using bioinformatics, high performance computing, and deep learning for biological advancements. Develop and accelerate bioinformatics software and alg...
nvidia.eightfold.ai
Great talk by @machine.learning.bio at the 5th Virtual @chembiotalks.bsky.social. He talked about the use of #MachineLearning and #BigData to address biological questions. Cool insights into both predicting functions and designing proteins ieeexplore.ieee.org/document/947... arxiv.org/abs/2503.00710
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning
Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new pred...
ieeexplore.ieee.org
GPU-accelerated MMseqs2 offers tremendous speedup for homology retrieval, protein structure prediction with ColabFold, and protein structure search with Foldseek. @martinsteinegger.bsky.social @milot.bsky.social @machine.learning.bio www.nature.com/articles/s41...
GPU-accelerated homology search with MMseqs2 - Nature Methods
Graphics processing unit-accelerated MMseqs2 offers tremendous speedups for homology retrieval from metagenomic databases, query-centered multiple sequence alignment generation for structure predictio...
nature.com
Looking forward to hearing about the potential of machine learning for #Biology and #DrugDiscovery from an industry perspective. Register for the Virtual @chembiotalks.bsky.social to hear the perspective of Chris Dallago (@machine.learning.bio) from Nvidia. #ChemBio #Chemsky #ML #MachineLearning
The third talk of the 5th Virtual ChemBioTalks will be given by Chris Dallago (@machine.learning.bio). He will talk about the research at NVIDIA into “Industrial BioML: the coming of age of machine learning for biology”. Make sure to register for free: https://cvent.me/G1geWW
(1/5) Venoms are a vast, largely untapped library of bioactive molecules—and our new paper in @natcomms.nature.com @natprot.nature.com reveals just how powerful they can be. 🐍⚡️
Computational exploration of global venoms for antimicrobial discovery with Venomics artificial intelligence - Nature Communications
Researchers used artificial intelligence to mine global venom proteomes and discovered novel peptides with antimicrobial activity. Several candidates showed efficacy against drug-resistant bacteria in...
nature.com
Excited to have participated in the 2025 Symposium on Generative AI in Molecule Discovery in beautiful Munich, along with amazing scientists and colleagues @machine.learning.bio, Francesca Grisoni, @ewaszczurek.bsky.social, @fabiantheis.bsky.social and more... 🔬🤖 events.hifis.net/event/2015/
Folddisco finds similar (dis)continuous 3D motifs in large protein structure databases. Its efficient index enables fast uncharacterized active site annotation, protein conformational state analysis and PPI interface comparison. 1/9🧶🧬 📄 www.biorxiv.org/content/10.1... 🌐 search.foldseek.com/folddisco
With contributions from fantastic colleagues @martinsteinegger.bsky.social , @mikeinouye.bsky.social, @jlistgarten.bsky.social , @ideasbyjin.bsky.social, @michael-heinzinger.bsky.social, and many more, the first CSHL volume on ML for Protein Science and Engineering is out: lnkd.in/dQdgGPpp
The final program is now online for the 5th Virtual @chembiotalks.bsky.social: web.cvent.com/event/60e9f3... Looking forward to talks by Sarah O'Connor, Chengqi Yi, @machine.learning.bio, @cathleenzeymer.bsky.social, @kellychibale.bsky.social, Jennifer Prescher and @craigmcrews.bsky.social.
The registration link and program for the 5th Virtual ChemBioTalks are finally available! If you are interested in the field of Chemical Biology, make sure to register for this free, virtual event on September 30th, 2025! cvent.me/G1geWW (1/4) #Chemsky #ChemBio
@michael-heinzinger.bsky.social and I are seeking talented postdocs to support for the Marie Skłodowska-Curie Fellowship! Join our international AI+biology team, collaborate on protein design, and access top labs in the US & EU. Interested? Apply by July 15! Details: machine.learning.bio/news/msca
Marie Skłodowska-Curie Fellowship
PostDoc Opportunity in AI for Protein Science Marie Skłodowska-Curie Fellowship
machine.learning.bio
Save the date for the Helmholtz Munich AI for Health Symposium 2025, devoted to the topic of Generative AI in Molecule Discovery! July 4, 2025 Helmholtz Munich Campus in Neuherberg, Germany What now? Register and submit abstracts (deadline May 2, 2025!) events.hifis.net/event/2015/r...
📢📢 "Proteina: Scaling Flow-based Protein Structure Generative Models" #ICLR2025 (Oral Presentation) 🔥 Project page: research.nvidia.com/labs/genair/... 📜 Paper: arxiv.org/abs/2503.00710 🛠️ Code and weights: github.com/NVIDIA-Digit... 🧵Details in thread... (1/n)
🔸Proteina is a fantastic collaboration with wonderful colleagues at NVIDIA: 🔥 Tomas Geffner*, @kdidi.bsky.social*, Zuobai Zhang*, Danny Reidenbach, Zhonglin Cao, @jyim.bsky.social , Mario Geiger, @machine.learning.bio, Emine Kucukbenli, @arashv.bsky.social, @karstenkreis.bsky.social* 🔥 (10/n)
Two major life updates: - I'm moving to Senior Applied Research Scientist in Digital Biology at NVIDIA (Jan '25) - I'm starting a new lab at Duke as Visiting Assistant Prof (early '25) Both roles focus on tackling hard problems in biological machine learning through collaborative research. Long 🧵
After a busy CASP/CAPRI year we resume our rolling CAPRI rounds, announcing the 1st target of 2025. It consists of an antibody-glycan complex. The glycan is heptyl α-D-mannopyranoside. Registration for this target is now open. pdbe.org/capri
The European Bioinformatics Institute < EMBL-EBI
EMBL-EBI
pdbe.org
We've a postdoc opening for our lab at the Broad: Cambridge MA! Must have experience in toxicology + data science Work on the wonderful OASIS dataset we are producing... Cell Painting, transcriptomics, proteomics in various liver cell and tissue models! broad.io/mlcbpostdoc
CDC just confirmed the first severe case of #H5N1 in the US in a patient in Louisiana. This virus seems to be the same genotype D1.1 that is spreading in birds at the moment (so not the cattle genotype B3.13) that severely sickened the teenager in Canada.
Protein language models excel at generating functional yet remarkably diverse artificial sequences. They however fail to naturally sample rare datapoints, like very high activities. In our new preprint, we show that RL can solve this without the need for additional data: arxiv.org/abs/2412.12979
Starting the new BSC IA factory. Financed by Europe, Spanish & Catalan governments + contributions from Portugal, Turkey, Romania New HPC and AI resources at the service of companies (hospitals included) With Juan Cruz Sec Estado Núria Montserrat Consellera Mateo Valero Dir @bsc-cns.bsky.social
We trained a model to co-generate protein sequence and structure by working in the ESMFold latent space, which encodes both. PLAID only requires sequences for training but generates all-atom structures! Really proud of @amyxlu.bsky.social 's effort leading this project end-to-end!
1/🧬 Excited to share PLAID, our new approach for co-generating sequence and all-atom protein structures by sampling from the latent space of ESMFold. This requires only sequences during training, which unlocks more data and annotations: bit.ly/plaid-proteins 🧵