Riboseek is a fast RNA/DNA search. More sensitive than nhmmer at 250x speed. Structure-aware realignment produces MSAs approaching rMSA quality. Plus 1.7M precomputed RNA MSAs, and an API to search your own 📄 www.biorxiv.org/content/10.6... 💾 github.com/steineggerla... 🌐 search.foldseek.com/riboseek
Milot Mirdita
@milot.bsky.social
Open source #bioinformatics at Sungkyunkwan University 🇰🇷 | former Steinegger Lab @ SNU, Söding Lab @ MPI-NAT | http://mstdn.science/@milotmirdita
Come work on exciting AI for biology projects in Paris!! 🇫🇷 Details and application below, deadline sept 1st!
🚨 JOB ALERT🚨 We are very excited to be hiring the lab's ✨very first postdoc✨! Work on new AI technologies for decoding antigen protein evolution in a fresh research environment, at the heart of Paris 🇫🇷 Details & link to apply: research.pasteur.fr/en/job/postd... Deadline: Sep 1st
Excited to share our now published paper @cp-cell.bsky.social highlighting advances in modeling the evolution of protein-protein interactions with MSA Pairformer. Big thanks to Zhidian Zhang, Olivia Tang, @eunbelivable.bsky.social @milot.bsky.social @martinsteinegger.bsky.social and @sokrypton.org!
Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer
Protein language models have excelled at modeling individual proteins, but extending these capabilities to protein complexes remains a major challenge. MSA Pairformer, a parameter-efficient protein la...
cell.com
Thanks for the advert, @martinsteinegger.bsky.social. If you're reading this and you're sitting on a pile of molecular dynamics simulations, please consider contributing them to MDRepo. This is the path to AI for dynamics (And if you're wondering: yes, there was only 1 person in the audience! 🤥)
@wheelerlab.org is talking about his effort to build a resource for protein dynamics (PDB for MD). MDRepo is a repository to store simulation data. This is really needed to push the needle in design, function and more. Please help make it successful by sharing your MD data. #ismb2026 🌐 mdrepo.org
New study by Alexander Henoch (@ahenoch.bsky.social), a PhD student in our group @hifmb.de and @awi.de, shows what it takes to bring gene synteny into microbial pangenomes, and what we learn about the variability landscape of genomes when we do that. See the pre-print here: doi.org/10.64898/202...
ICML is happening in Seoul this year, and I’ve been getting several messages about lab visits. Who else will be in town and would like to meet? @milot.bsky.social lab and mine are planning a dinner on July 7th, after the reception. Let me know if you’re interested!
🧬 New preprint! We clustered 5.6 million bacterial genomes into genomically cohesive units (GCUs) 500× faster than existing tools. (In just 14 hours, 16.5 GB RAM using 48 CPUs). 🦠🐙Meet gemsparcl 💎✨! www.biorxiv.org/content/10.6...
Does your designed active site already exist in nature? Is an uncharacterized protein hiding a catalytic site or a pocket? Folddisco answers both, searching millions of structures for a 3D motif in seconds. @natbiotech.nature.com 🧬 📄 www.nature.com/articles/s41... 🧵1/7👇
Structural motif search across the protein universe with Folddisco - Nature Biotechnology
Folddisco enables protein structural motif search in million scale databases.
nature.com
10 years after the first FAMSA paper, its successor is now published in Nat Biotech! We believe that FAMSA2 can enable analyses of large protein collections that were previously unattainable. Thank you, Andrzej and Cedric, for great collaboration www.nature.com/articles/s41...
Fast and accurate multiple-protein-sequence alignment at scale with FAMSA2 - Nature Biotechnology
FAMSA2 accurately aligns millions of protein sequences at high speed.
nature.com
Metabuli & Metabuli App v1.2 improve novel species classification with higher precision and recall. New light mode is 1.8× faster and requires 50% less storage while keeping precision. New RefSeq, GTDB, HRGM, and HROM databases added. 💾 github.com/steineggerla... 📄 doi.org/10.64898/2026.03.13.711249
Whenever I presented Phold, I was frequently asked "can you do the same beyond phages?" We ( @oschwengers.bsky.social @linsalrob.bsky.social @binomicalabs.org et al) finally did it with Baktfold github.com/gbouras13/ba... www.biorxiv.org/content/10.6...
GitHub - gbouras13/baktfold: Rapid & standardized genome annotation using protein structural information
Rapid & standardized genome annotation using protein structural information - gbouras13/baktfold
github.com
45 novel protein folds in the updated AFESM (AFDB + ESMatlas) manuscript: • 12 high-confidence folds in AFESM • 33 by ColabFold-repredicting 2.3M low-quality domains We show AFDB captures most domains already and ESMfold struggles with novelty 🌏 afesm.foldseek.com 📄 biorxiv.org/content/10.1...
AFESM Clusters
Foldseek clustered 820M AlphaFold DB + ESMatlas structures
afesm.foldseek.com
Baktfold: Sensitive protein functional annotation across the microbial tree of life using structural information https://www.biorxiv.org/content/10.64898/2026.03.31.715528v1
the web application is available at : https://brigx.genomicx.org/ I would be interested to hear feedback regarding bugs or any unexpected behaviour, any suggestions for the user interface, or any features you feel would be useful .
BRIGX - Browser-Based Ring Image Generator
Circular comparative genome visualization tool running entirely in your browser
brigx.genomicx.org
Clustering proteins using DIAMOND is out now @natmethods.nature.com www.nature.com/articles/s41...
Clustering the protein universe of life using DIAMOND DeepClust - Nature Methods
DIAMOND DeepClust provides an ultra-fast clustering method for organizing the protein universe of life at low sequence identity, enabling large-scale dimensionality reduction and improving downstream ...
nature.com
How much protein diversity can Life on Earth actually generate? With DIAMOND DeepClust, we show how billions of proteins across the tree of life can be clustered at low-identity for downstream analytics tasks. 📚Paper: www.nature.com/articles/s41... 💻Code: github.com/bbuchfink/di...
My group at MIT is seeking a research scientist with a strong *experimental* background to lead and help shape the lab’s experimental infrastructure, supporting efforts to advance AI-driven enzyme discovery and characterization. See the full JD here: acrobat.adobe.com/id/urn:aaid:...
Adobe Acrobat
acrobat.adobe.com
Meet evedesign: open-source AI, accessible protein design ✅Combine models for multiobjective optimization ✅Integrate experimental data ✅ Run on your own infrastructure 📄Paper: www.biorxiv.org/content/10.6... 💻Code: github.com/evedesignbio 🌐Webserver: evedesign.bio Collaborate: hello@evedesign.bio
evedesign: accessible biosequence design with a unified framework
Unified protein design for computational researchers and experimentalists
deboramarkslab.substack.com
@ecallaway.bsky.social wrote a news article on our AlphaFold complex work. Thank you for covering it. 📄 www.nature.com/articles/d41...
AlphaFold hits ‘next level’: the AI database now includes protein pairing
The database of 200 million protein-structure predictions now includes homodimers, adding new biological relevance.
nature.com
AlphaFold database has entered the era of complexes. Together with NVIDIA, DeepMind and EBI, we use ColabFold, OpenFold and MMseqs2-GPU to predict ~31 million complexes (homo & hetro-dimers) resulting in 1.8 million high-quality predictions 📄 research.nvidia.com/labs/dbr/ass... 🌐 alphafold.ebi.ac.uk
You asked, we listened. Millions of AI-predicted protein complex structures are now available in the #AlphaFold Database. This spans homodimers from 20 of the most studied species, including humans, as well as the World Health Organization’s priority pathogens list. www.ebi.ac.uk/about/news/t...
Efficient protein structure prediction fromcompact computers to datacenters withOpenFold-TRT https://www.biorxiv.org/content/10.64898/2026.03.11.711233v1
ProteinTTT is now easy to run on Hugging Face Spaces and Google Colab. We’ll also be presenting the paper at ICLR 2026 🇧🇷 🤗 Hugging Face Space: huggingface.co/spaces/pimen... ⚙️ Google Colab: colab.research.google.com/drive/1l_h7c... 🧵👇
My first manuscript in MPI colours! With @tothpetroczylab.bsky.social, we show that AlphaFold PAE-derived contact probabilities are well calibrated to the fraction of true interface contacts across experimentally determined protein dimers. www.biorxiv.org/content/10.6...
Can't wait to release a 10-year-old birthday version for SeqKit! - 10 years - 2 papers, 3500 citations - 20 contributors - 40 subcommands - 880 commits - 500 issues - 685.5K Bioconda total downloads Thank you all, dear contributors and users! I'll keep maintaining it. github.com/shenwei356/s...
Release SeqKit v2.13.0 (10-year-old birthday version) · shenwei356/seqkit
Changelog SeqKit is 10 years old! SeqKit v2.13.0 - 2026-02-28 seqkit: add support for reading and writing LZ4 compression format. new command: seqkit sample2: improved seqkit sample by @stahiga....
github.com
At the 132nd Internat. Titisee Conference on Biology 2.0: The AI Revolution in Biology & Medicine From sequence→function models 🧬 to protein & generative structure models 🧪 to AI of cell states & perturbations 🧫 Great science, great friends, beautiful lake. Thanks @BIFonds!
New version of our preprint on bioRxiv about bioRxiv up. Now that’s what I call a revision – 6 years after the first version! It has new data about our progress and highlights from a massive user survey. 1/n www.biorxiv.org/content/10.1...
biorxiv.org
Can we simulate realistic evolutionary trajectories and “replay the tape of life”? In this work, we propose a flexible, generalizable deep learning framework for modeling how the entire protein sequence evolves over time while capturing complex interactions across sites. 1/n doi.org/10.64898/202...
doi.org
Our new review on genome annotation just appeared in @naturerevgenet.bsky.social, with a particular focus on the human genome, with Hayden Ji and Mihaela Pertea: rdcu.be/e4mI1
Annotating genomes at increased scale and resolution
Nature Reviews Genetics - In this Review, Ji et al. overview how rapidly advancing experimental and computational methods are enabling improved and automated annotation of gene structure and...
rdcu.be