10 years ago today @casey.greenelab.com launched our review "Opportunities and obstacles for deep learning in biology and medicine" greenelab.github.io/deep-review/. It was a widely collaborative project written on GitHub, which lead to the creation of manubot.org.
Anthony Gitter
@anthonygitter.bsky.social
Computational biologist; Associate Prof. at University of Wisconsin-Madison; Jeanne M. Rowe Chair at Morgridge Institute https://gitterlab.org/
Thanks for the advert, @martinsteinegger.bsky.social. If you're reading this and you're sitting on a pile of molecular dynamics simulations, please consider contributing them to MDRepo. This is the path to AI for dynamics (And if you're wondering: yes, there was only 1 person in the audience! 🤥)
@wheelerlab.org is talking about his effort to build a resource for protein dynamics (PDB for MD). MDRepo is a repository to store simulation data. This is really needed to push the needle in design, function and more. Please help make it successful by sharing your MD data. #ismb2026 🌐 mdrepo.org
The @qedscience.bsky.social "impact" score generated a lot of discussion on ranking preprints, including ideas on multi-dimensional rankings that are user specific. Besides describing the ideas, we can now prototype them (with Claude in this case). Here is Salient salient-sgeu6fs5ua-uc.a.run.app
Rising — Salient
salient-sgeu6fs5ua-uc.a.run.app
A great story stemming from collaboration between @anthonygitter.bsky.social and Nate Wlodarchak, now in Colorado at the Rocky Mountain Regional VA Medical Center. ⬇️
Former Cancer Biology graduate student Nate Wlodarchak is collaborating with @morgridgeinstitute.bsky.social scientists to find new drugs to treat tuberculosis. Read about it here: morgridge.org/story/unlock...
Genomic and other biological data are in scope for this data scientist position if anyone in that area is looking.
The Data Science Institute is looking for data scientists who thrive in a research environment to join our team! Learn more and apply by July 11: jobs.wisc.edu/jobs/data-sc...
In the W-S lab's first preprint, we describe how genomic language models know something about RNA thermodynamics. Though we think this is cool, things get tricky! A growing practice for interpreting LMs is to perturb input tokens, often called "Categorical Jacobian": 👇
Now published in NSMB! Paper: doi.org/10.1038/s415... Full PDF: rdcu.be/fhBtI Overview of additions since the preprint👇 (1/5)
Evaluating generalization in protein–ligand cofolding methods - Nature Structural & Molecular Biology
This work introduces the Runs N’ Poses dataset for benchmarking deep learning methods on the protein–ligand complex prediction task. It shows that current methods rely on memorization, challenging the...
doi.org
Excited to share our latest preprint evaluating AlphaFold3, Boltz-1, Chai-1 and Protenix for predicting protein-ligand interactions, featuring our newly introduced benchmark dataset 🌹Runs N’ Poses🌹! www.biorxiv.org/content/10.1... 🧵👇 (1/n)
I am happy to share a review I recently wrote on the design of peptide binders. It gives an overview of experimentally validated tools and discusses the challenges of why peptide design is more difficult than the design of classical protein binders. www.chimia.ch/chimia/artic...
Fantastic analysis from the OpenADMET team (Maria Castellanos, Hugo MacDermott-Opeskin) showing that the zero-shot ADMET models ADMETlab 3.0 and ADMET-AI generalize poorly to their recent OpenADMET-ExpansionRx Blind Challenge data openadmet.ghost.io/zero-shot-ex...
Lessons Learned from the OpenADMET-ExpansionRx Blind Challenge: Can We Trust Zero-Shot ADMET Predictions?
Maria Castellanos Hugo MacDermott-Opeskin It’s been more than a month since the OpenADMET-ExpansionRx challenge wrapped up, but the conversation is just getting started. Launched on October 27, 2025...
openadmet.ghost.io
Is #AI hitting a plateau in structure prediction? Help us find out at CASP17! 🧪🧬 Calling for Targets: Immune Complexes, protein - ligand complexes, RNA/DNA, conformational ensembles, membrane proteins, viral origins, and large complexes. The Rule of Thumb: If AF3 can’t model it, we want it.
We have started a project trying to predic the interactions/structures of all yeast protein pairs using an AlphaFold pooling approach. We are making the current dataset open and we welcome collaborations. www.evocellnet.com/2026/03/mapp...
Mapping the yeast atructural interactome with AlphaFold3: an open call for collaboration
We are excited to announce the early-stage release of our S. cerevisiae structural interactome mapping project. Using AlphaFold3 (AF3), w...
evocellnet.com
Can we simulate realistic evolutionary trajectories and “replay the tape of life”? In this work, we propose a flexible, generalizable deep learning framework for modeling how the entire protein sequence evolves over time while capturing complex interactions across sites. 1/n doi.org/10.64898/202...
doi.org
Can proteins fold and function with half of the amino acid alphabet? Using only 10 residues, we designed stable, mutation-resilient structures—no aromatics or basics involved. A minimalist foundation for ancient biology and synthetic design. tinyurl.com/37t8br4v #ProteinDesign #OriginsOfLife
Ancient amino acid sets enable stable protein folds
Early proteins likely arose from a chemically limited set of amino acids available through prebiotic chemistry, raising a central question in molecular evolution: could such primitive compositions yie...
tinyurl.com
My time in @martinsteinegger.bsky.social's group is ending, but I’m staying in Korea to build a lab at Sungkyunkwan University School of Medicine. If you or someone you know is interested in molecular machine learning and open-source bioinformatics, please reach out. I am hiring! mirdita.org
Mirdita Lab - Laboratory for Computational Biology & Molecular Machine Learning
Mirdita Lab builds scalable bioinformatics methods.
mirdita.org
I'm really excited to break up the holiday relaxation time with a new preprint that benchmarks AlphaFold3 (AF3)/“co-folding” methods with 2 new stringent performance tests. Thread below - but first some links: A longer take: fraserlab.com/2025/12/29/k... Preprint: www.biorxiv.org/content/10.6...
Know when to co-fold'em
This is the official web page for the James Fraser Lab at UCSF.
fraserlab.com
New preprint🚨 Imagine (re)designing a protein via inverse folding. AF2 predicts the designed sequence to a structure with pLDDT 94 & you get 1.8 Å RMSD to the input. Perfect design? What if I told u that the structure has 4 solvent-exposed Trp and 3 Pro where a Gly should be? Why to be wary🧵👇
Excited for our new paper on a genome language model for viruses in @natcomms.nature.com: "Protein Set Transformer: a protein-based genome language model to power high-diversity viromics"! Led by PhD student Cody Martin in collaboration with @anthonygitter.bsky.social doi.org/10.1038/s414...
Protein Set Transformer: a protein-based genome language model to power high-diversity viromics - Nature Communications
A genome language model, Protein Set Transformer, trained on viral datasets, uncovers evolutionary rules of protein content and organization driving precise virus identification, host prediction, and ...
doi.org
What are good places to post an unsolicited manuscript peer review these days? I don't have a blog. I read manuscripts across arXiv, bioRxiv, ChemRxiv, OpenReview, random white papers, journals, etc. Do I dump it on Zenodo, post it here, and send it to the authors?
Our Assay2Mol manuscript was published at EMNLP 2025 doi.org/10.18653/v1/... See the preprint thread below for a summary of the methodology, results, and code. We added more control experiments in this version related to protein sequence identity and generated molecule size.
Our preprint Assay2Mol introduces uses PubChem chemical screening data as context when generating molecules with large language models. It uses assay descriptions and protocols to find relevant assays and that text plus active/inactive molecules as context for generation. 1/
@hkws.bsky.social and I are creating the Madison AI for Proteins (MAIP) group to discuss early-stage research at monthly meetups, share computational resources, and grow this local community. Visit mad-ai-proteins.github.io to sign up for announcements and watch for our 2026 events.
MAIP
Madison AI for Proteins
mad-ai-proteins.github.io
Something fun and sciencey is coming soon to Madison
Something fun and sciencey is coming soon to Madison
The journal version of our Multi-omic Pathway Analysis of Cells (MPAC) software is now out: doi.org/10.1093/bioi... MPAC uses biological pathway graphs to model DNA copy number and gene expression changes and infer activity states of all pathway members.
AI + physics for protein engineering 🚀 Our collaboration with @anthonygitter.bsky.social is out in Nature Methods! We use synthetic data from molecular modeling to pretrain protein language models. Congrats to Sam Gelman and the team! 🔗 www.nature.com/articles/s41...
Biophysics-based protein language models for protein engineering - Nature Methods
Mutational effect transfer learning (METL) is a protein language model framework that unites machine learning and biophysical modeling. Transformer-based neural networks are pretrained on biophysical simulation data to capture fundamental relationships between protein sequence, structure and energetics.
nature.com
Does anyone know whether there's a functioning API to ESMfold? (api.esmatlas.com/foldSequence... gives me Service Temporarily Unavailable)
The journal version of "Biophysics-based protein language models for protein engineering" with @philromero.bsky.social is live! Mutational Effect Transfer Learning (METL) is a protein language model trained on biophysical simulations that we use for protein engineering. 1/ doi.org/10.1038/s415...
Biophysics-based protein language models for protein engineering - Nature Methods
Mutational effect transfer learning (METL) is a protein language model framework that unites machine learning and biophysical modeling. Transformer-based neural networks are pretrained on biophysical ...
doi.org
The journal version of our paper 'Chemical Language Model Linker: Blending Text and Molecules with Modular Adapters' is out doi.org/10.1021/acs.... ChemLML is a method for text-based conditional molecule generation that uses pretrained text models like SciBERT, Galactica, or T5.
Chemical Language Model Linker: Blending Text and Molecules with Modular Adapters
The development of large language models and multimodal models has enabled the appealing idea of generating novel molecules from text descriptions. Generative modeling would shift the paradigm from relying on large-scale chemical screening to find molecules with desired properties to directly generating those molecules. However, multimodal models combining text and molecules are often trained from scratch, without leveraging existing high-quality pretrained models. Training from scratch consumes more computational resources and prohibits model scaling. In contrast, we propose a lightweight adapter-based strategy named Chemical Language Model Linker (ChemLML). ChemLML blends the two single domain models and obtains conditional molecular generation from text descriptions while still operating in the specialized embedding spaces of the molecular domain. ChemLML can tailor diverse pretrained text models for molecule generation by training relatively few adapter parameters. We find that the choice of molecular representation used within ChemLML, SMILES versus SELFIES, has a strong influence on conditional molecular generation performance. SMILES is often preferable despite not guaranteeing valid molecules. We raise issues in using the entire PubChem data set of molecules and their associated descriptions for evaluating molecule generation and provide a filtered version of the data set as a generation test set. To demonstrate how ChemLML could be used in practice, we generate candidate protein inhibitors and use docking to assess their quality and also generate candidate membrane permeable molecules.
doi.org
🚨New paper 🚨 Can protein language models help us fight viral outbreaks? Not yet. Here’s why 🧵👇 1/12
Our preprint Assay2Mol introduces uses PubChem chemical screening data as context when generating molecules with large language models. It uses assay descriptions and protocols to find relevant assays and that text plus active/inactive molecules as context for generation. 1/