Metabuli & Metabuli App v1.2 improve novel species classification with higher precision and recall. New light mode is 1.8× faster and requires 50% less storage while keeping precision. New RefSeq, GTDB, HRGM, and HROM databases added. 💾 github.com/steineggerla... 📄 doi.org/10.64898/2026.03.13.711249
Dongwook Kim
@dongwookkim.bsky.social
Developing fast and easy methods for #phylogenetics and #bioinformatics | PhD in Bioinformatics | Postdoc @ Comparative Genomics Lab, UNIL/SIB🇨🇭| Formerly @ Steinegger Lab, SNU🇰🇷 | he/him
I am pleased to share that our paper is now published in Cell! www.cell.com/cell/fulltex... I am deeply grateful to all co-authors for making this possible. This work was made possible through the guidance of Dr. Peer Bork. I share this in grateful memory and with deep respect for his mentorship.
Planetary microbiome structure and generalist-driven gene flow across disparate habitats
A planetary-scale analysis of over 85,000 metagenomes establishes a framework for exploring the structure and drivers of global microbial habitats, revealing that generalist species bridge ecological ...
cell.com
In the largest study of its kind, scientists in the Bork Group at EMBL have found that a small subset of microbes can carry and transfer genes across disparate habitats, creating a planet-wide, interconnected network of microbiomes 🌍 🦠 🔗 Read more here: www.embl.org/news/science...
FoldMason is out now in @science.org. It generates accurate multiple structure alignments for thousands of protein structures in seconds. Great work by Cameron L. M. Gilchrist and @milot.bsky.social. 📄 www.science.org/doi/10.1126/... 🌐 search.foldseek.com/foldmason 💾 github.com/steineggerla...
Multiple protein structure alignment at scale with FoldMason
Protein structure is conserved beyond sequence, making multiple structural alignment (MSTA) essential for analyzing distantly related proteins. Computational prediction methods have vastly extended ou...
science.org
Can ever-increasing sequence databases improve phylogenetic reconstruction of a gene family? Our new preprint introduces AmpliPhy, a pipeline that automates homolog enrichment to improve gene tree inference, built on a robust phylogenomic benchmark scheme. 🧵1/n 📃 doi.org/10.64898/2026.01.26.701724
AmpliPhy improves gene trees by adding homologs without affecting alignments
In phylogenomics, gene tree reconstruction depends on multiple sequence alignment (MSA) and tree inference, and ongoing work continues to improve inference quality. Denser taxon sampling has been associated with improved gene tree inference, suggesting that adding homologs could be a practical route to higher accuracy as sequence databases continue to expand. However, adding sequences can influence multiple steps of typical inference pipelines, and little is known on its specific effect on the multiple sequence alignment, tree reconstruction, and rooting steps. We performed a large-scale empirical benchmark to quantify how homolog enrichment affects alignment and phylogenetic inference. Using an enrichment-impoverishment design and a measure of tree accuracy based on taxonomic congruence, we found that enrichment consistently improves tree inference quality, while effects on alignment quality are marginal. We show that this improvement is associated with accurate root placement on enriched trees when sensitive homolog search is accompanied. Notably, much of the benefit can be retained with relatively compact alignments produced by sequence addition. Building on these observations, we provide a tool, AmpliPhy, which efficiently improves phylogenetic reconstruction of protein families through homolog enrichment. The AmpliPhy open-source pipeline software is available at https://github.com/DessimozLab/ampliphy. ### Competing Interest Statement The authors have declared no competing interest. Swiss National Science Foundation, https://ror.org/00yjd3n13, 216623, 10005715
doi.org
My time in @martinsteinegger.bsky.social's group is ending, but I’m staying in Korea to build a lab at Sungkyunkwan University School of Medicine. If you or someone you know is interested in molecular machine learning and open-source bioinformatics, please reach out. I am hiring! mirdita.org
Mirdita Lab - Laboratory for Computational Biology & Molecular Machine Learning
Mirdita Lab builds scalable bioinformatics methods.
mirdita.org
Stoked to finally have a preprint out for Phold, our tool that uses protein structural information to enhance phage genome annotation #phagesky 1/n www.biorxiv.org/content/10.1...
Protein Structure Informed Bacteriophage Genome Annotation with Phold
Bacteriophage (phage) genome annotation is essential for understanding their functional potential and suitability for use as therapeutic agents. Here we introduce Phold, an annotation framework utilis...
biorxiv.org
Our new preprint is out! www.biorxiv.org/content/10.1... In this study, we present the largest systematic analysis of microbiome structure and function, integrating 85K uniformly processed metagenomes from diverse habitats worldwide. @podlesny.bsky.social @jonas-bio.bsky.social @borklab.bsky.social
Planetary microbiome structure and generalist-driven gene flow across disparate habitats
Microbes are ubiquitous on Earth, forming microbiomes that sustain macroscopic life and biogeochemical cycles. Microbial dispersion, driven by natural processes and human activities, interconnects mic...
biorxiv.org
OrthoFinder just dropped a major update It’s faster, more accurate, and ready for thousands of genomes Let’s break it down (1/10) github.com/OrthoFinder/... www.biorxiv.org/content/10.1...
Folddisco finds similar (dis)continuous 3D motifs in large protein structure databases. Its efficient index enables fast uncharacterized active site annotation, protein conformational state analysis and PPI interface comparison. 1/9🧶🧬 📄 www.biorxiv.org/content/10.1... 🌐 search.foldseek.com/folddisco
New paper from the lab from Sriram Garg in my group. We introduce a general substitution matrix for structural phylogenetics. I think this is a big deal, so read on below if you think deep history is important. academic.oup.com/mbe/advance-...
A general substitution matrix for structural phylogenetics.
Abstract. Sequence-based maximum likelihood (ML) phylogenetics is a widely used method for inferring evolutionary relationships, which has illuminated the
academic.oup.com
Unicore is now published on GBE 🚀 Unicore rapidly identifies structural single-copy core genes from input species proteomes for phylogenetic analysis. Powered by Foldseek and ProstT5, Unicore enables linear-scale structure-based phylogeny of any given set of taxa. 🧵1/n 📃 doi.org/10.1093/gbe/evaf109
AFESM: a metagenomic guide through the protein structure universe! We clustered 821M structures (AFDB&ESMatlas) into 5.12M groups; revealing biome-specific groups, only 1 new fold even after AlphaFold2 re-prediction & many novel domain combos. 🧵 🌐 afesm.foldseek.com 📄 www.biorxiv.org/content/10.1...
Visit our posters at #RECOMB2025 for: Structural: MSAs, Virus DB, Core Genes, Motif Discovery, Multimer Clustering & Search, pLM Foldseek, Environmental analysis Metagenomics: Classification & Metabuli App GPU-based & RNA search, Proteome clustering, Novel Ribozyme discovery & get Marv stickers!
Not really my announcement to make--I am but a lesser co-author--but IQ-TREE 3 has just been released! (Most credit to Minh Bui and @roblanfear.bsky.social and their labs) ecoevorxiv.org/repository/v...
IQ-TREE 3: Phylogenomic Inference Software using Complex Evolutionary Models
ecoevorxiv.org
🚀 #AlphaFold Database update AlphaFold DB now integrates The Encyclopedia of Domains (TED) – a resource designed to systematically identify & classify structural domains within AlphaFold-predicted protein structures. www.ebi.ac.uk/about/news/u... @pdbeurope.bsky.social
The PAN-GO paper is a remarkable milestone. It not only provides the most comprehensive picture of human gene function to date, but also carefully maps this knowledge across the tree of life! Congratulations @marcfeuermann.bsky.social, Pascale Gaudet & collaborators! www.sib.swiss/news/sib-hel...
SIB helps create most complete, accurate resource for human gene functions
For the first time, biodata from humans have been integrated with that of other organisms to provide the most comprehensive picture of human gene function to date. The new ‘PAN-GO’ resource used e...
sib.swiss
Nature research paper: A compendium of human gene functions derived from evolutionary modelling https://go.nature.com/4gXxIRg
In our latest review, we explore 12 deep-learning tools for metagenomic analysis, covering their strengths, limitations, and key applications. We hope it serves as both a resource and inspiration for new ways to analyze metagenomic data. Great work by Eli Levy Karin! 📄 doi.org/10.1093/nsr/...
FastOMA is out now in Nature Methods 🎉: nature.com/articles/s41592-024-02552-8 A new orthology inference algorithm that scales linearly and is highly accurate. FastOMA can process all >2000 eukaryotic UniProt ref proteomes <24 hours 🚀. Try it out github.com/DessimozLab/fastoma @dessimoz.bsky.social
Unicore identifies single-copy protein structures across genomes using Foldseek, bypassing slow structure predictions by utilizing 3Di predictions from ProstT5, enabling rapid phylogenetic inference at the tree-of-life scale. 1/n 📄 www.biorxiv.org/content/10.1... 💾 github.com/steineggerla...
Unicore enables scalable and accurate phylogenetic reconstruction with structural core genes https://www.biorxiv.org/content/10.1101/2024.12.22.629535v1
Scientists, academics, researchers: We’re excited to share that @altmetric.com is now tracking mentions of your research on Bluesky! 🧪
There are already many articles for which there is more attention on Bluesky than on other comparable micro-blogging sites, meaning the academic community and the general public have clearly adopted Bluesky as one of its core places to disseminate and discuss new research. A Place of Joy.
South Korean citizens helped lawmakers scale the National Assembly walls so they could bypass military barricades and vote against martial law.
Reminder for newcomers that bioRxiv has Bluesky accounts in every subject category - great way to keep up (please re-skeet) connect.biorxiv.org/news/2023/09...
bioRxiv expands on Mastodon and Bluesky
bioRxiv - the preprint server for biology, operated by Cold Spring Harbor Laboratory, a research and educational institution
connect.biorxiv.org
Interested in bioinformatics method development for proteins, structures or metagenomic analysis? Please check out my lab’s starter pack! 🔗 go.bsky.app/VJhXcSs
MMseqs2 Release 16 Highlights: GPU-accelerated search📄, ORF or new 6-frame translated search modes, contig taxonomy always keeps the longest ORF, bug fixes (reduced memory and higher sensitivity) and relicensed as MIT 📄 biorxiv.org/content/10.1... 💾 mmseqs.com and 🐍Bioconda 🖥️🧬🧶
What did the Last Eukaryotic Common Ancestor (#LECA) look like? Consensus View in #PLOSBiology; massive authorship including @AncestralState, @lauraeme.bsky.social, John Archbald, @andrewjroger.bsky.social, @dackslabecb.bsky.social, Jeremy Wideman. plos.io/4g0alq4
Our Big Fantastic Virus Database (BFVD) is now published NAR! It contains protein structure predictions of major viral clades, enhanced by petabase-scale homology search and it's explorable on the web. 🌐 bfvd.foldseek.com 💾 bfvd.steineggerlab.workers.dev 📄 academic.oup.com/nar/advance-...