Riboseek is a fast RNA/DNA search. More sensitive than nhmmer at 250x speed. Structure-aware realignment produces MSAs approaching rMSA quality. Plus 1.7M precomputed RNA MSAs, and an API to search your own 📄 www.biorxiv.org/content/10.6... 💾 github.com/steineggerla... 🌐 search.foldseek.com/riboseek
Gu Zhenhao
@guzhenhao.bsky.social
PhD student at NUS Computing / Genome Institute of Singapore. Alto clef enjoyer. My playlist: https://www.youtube.com/playlist?list=PLPgSTwKT0Yv8bm859AW-s3lqjEKPBCS0h
New preprint🚨(*long* time coming)! We introduce seqproc: an efficient, flexible, and concise tool for describing and transforming sequencing-read geometry. The aim: make complex protocols easy to specify & transform without giving up speed or accuracy. www.biorxiv.org/content/10.6... 🧵
Significant update to the AllTheBacteria paper, including discovering new antimicrobial peptides and testing in vitro and vivo. This has grown into a fantastic collaboration!
Biology has plenty of data—the challenge is making it usable. AllTheBacteria transforms 2.44 million public bacterial and archaeal genomes into an open, uniformly processed, searchable, AI-ready resource. www.biorxiv.org/content/10.1...
New Nature Reviews Genetics paper out! How are graph-based pangenomes removing reference bias and opening up new possibilities in GWAS and rare- and common disease genetics? 🔗 nature.com/articles/s41576-026-00987-7 #pangenome #geneticdiversity #referencebias
Building and applying pangenome references to capture genetic diversity - Nature Reviews Genetics
Pangenomes are genome references that integrate sequences from multiple individuals into graph-based or multi-haplotype representations, capturing genetic variation beyond a single linear reference. H...
nature.com
1/ Excited to share the newest tool in the pangenome MUMiverse: Shredtools! Shredtools enables a user to navigate the pangenome coordinate system with multi-MUMs. More in the thread🧵 Code: github.com/vikshiv/shredtools Interactive tool for querying HPRC assemblies: vikshiv.github.io/shredtools
Navigating the pangenome coordinate system with Shredtools
Existing notions of pangenome coordinates rely on hard-to-compute multiple sequence alignments. On the other hand, pangenome-wide exact unique matches (multi-MUMs) can be computed efficiently, and rep...
biorxiv.org
Our #ISMB2026 paper is now online! I’m excited to present it at HiTSeq on July 13. Many thanks to my advisor @pashadag.bsky.social. I’m also grateful to @iscb.bsky.social for awarding me the Conference Fellowship. If you’ll be at ISMB, feel free to stop by my talk and say hi. I’d love to connect!
The gift of novelty: repeat-robust k-mer-based estimators of mutation rates
AbstractMotivation. Estimating mutation rates between evolutionarily related sequences is a central problem in molecular evolution. Due to the rapid expans
academic.oup.com
Ever wanted to quickly check host content of DNA sequences? bede.im/sapiometer
It's my pleasure to introduce for the first time r-news.net, a project I had in my head for quite a long time (since my PhD!). With AI improving literature review quite substantially, we still miss a critical step: What does the community as a whole like the most?
Top Stories — RNews
r-news.net
So proud to see this study out — a real honor to be involved in work from my PhD lab!
High-quality phage assembly from metagenomes with PALACE www.nature.com/articles/s41...
Movi 2 has appeared (as an advance article) in Bioinformatics 🧬 Faster, leaner pangenome queries — half the memory of Movi 1, ~30% faster. Paper: academic.oup.com/bioinformati... Code: github.com/mohsenzakeri/Movi (1/6)
Validate User
academic.oup.com
Minibwa is a hybrid of bwa-mem and minimap2 and the successor of bwa-mem for short-read mapping. ~4X/2.5X as fast as bwa-mem/bwa-mem2 for WGS reads at comparable accuracy. Native support of directional bisulfite-seq. Applicable to long reads. Preprint at arxiv.org/abs/2606.15357
Just submitted my PhD thesis on algorithms for fast, large-scale k-mer-based sequence analysis. It's now available to read at phd.martayan.org Take a look and feel free to share! #Bioinformatics #PhDone
Algorithm design and implementation for the scale of sequencing data
phd.martayan.org
Nanopore sequencing provides not just long reads, but the the raw signal data can also be used to identify RNA and DNA modifications. This repository (and the associated review) lists some of the great tools that have been developed www.cell.com/trends/genet...
Beyond sequencing: machine learning algorithms extract biology hidden in Nanopore signal data
Nanopore sequencing provides signal data corresponding to the nucleotide motifs sequenced. Through machine learning-based methods, these signals are translated into long-read sequences that overcome t...
cell.com
We’ve updated our awesome-nanopore list! The list is a community-curated list of ONT software tools, please feel free to check it out & contribute: github.com/GoekeLab/awesome-nanopore #Nanopore #Bioinformatics #LongReadSequencing
Does your designed active site already exist in nature? Is an uncharacterized protein hiding a catalytic site or a pocket? Folddisco answers both, searching millions of structures for a 3D motif in seconds. @natbiotech.nature.com 🧬 📄 www.nature.com/articles/s41... 🧵1/7👇
Structural motif search across the protein universe with Folddisco - Nature Biotechnology
Folddisco enables protein structural motif search in million scale databases.
nature.com
Glad to announce that our new long-read metagenomic SNP caller, SNooPy, is published in NAR! Existing long-read SNP callers are not designed for metagenomic data, check out our new solution 👉https://academic.oup.com/nar/article/54/10/gkag556/8700491 @narjournal.bsky.social
Our method, savont, for generating amplicon sequence variants (ASVs) for long-read amplicons is now on bioRxiv. Work with @lh3lh3.bsky.social and help from @mkddueholm.bsky.social and team (Marie Riisgaard-Jensen, @kirk3gaard.bsky.social, Kasper Skytte Andersen) github.com/bluenote-157... 1/6
Sensitive long-read amplicon sequence variant recovery with savont https://www.biorxiv.org/content/10.64898/2026.05.26.727271v1
Introducing nail - a Rust implementation of profile HMM sequence alignment for proteins. Near-HMMER sensitivity, but a lot faster: www.biorxiv.org/content/10.1... github.com/TravisWheele...
biorxiv.org
Revamped the teaching materials site. Light/dark mode switch at the top. More consistent, compact presentation of all the materials. + Icons and such! www.langmead-lab.org/teaching.html
KaryoScope: rapid, alignment-free sequence annotation for the pangenome era https://www.biorxiv.org/content/10.64898/2026.05.15.725544v1
Population-level structural variant characterization using pangenome graphs www.nature.com/articles/s41... Swave introduces 'projection waves` to summarize the dotplot images, capturing mapping patterns in pangenomes. A recurrent neural networks distinguishes true SV from background noise/repeats.
🔍 New paper in Bioinformatics Advances: "Kaminari: A frugal colored index for approximate k-mer queries" Read it here: https://doi.org/10.1093/bioadv/vbag120 Authors include: @yhhshb.bsky.social, @yoann.bsky.social, @robp.bsky.social, @pierrepeterlongo.bsky.social, @jermp.bsky.social
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
doi.org
1/4 Hash functions in genomic sequence analysis (tinyurl.com/4kk9ccmt) : a new survey written together with Ke Chen, Xiang Li, Qian Shi, and Mingfu Shao. Before submitting it, we are posting it online to get feedback from the community.
HERRO has been published in @nature.com nature.com/articles/s41.... This achievement is a result of the great work by Dominik Stanojevic, with contributions from Dehui Lin, @sergeynurk.bsky.social, and Paola Florez de Sessions Welcome to the era of high-quality genome assemblies supported by AI.
Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads - Nature
Nature - Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads
nature.com
Excited to finally share our new preprint on bioRxiv describing Verticall (github.com/rrwick/Verti...), a robust & efficient tool for building recombination-free bacterial phylogenies. Huge thanks to @rrwick.bsky.social & @katholt.bsky.social for this incredible work! www.biorxiv.org/content/10.6...
biorxiv.org
QCatch is now published in Bioinformatics (academic.oup.com/bioinformati...)! Great work from Yuan and Dongze for quality control and analysis downstream of simpleaf/alevin-fry (taking advantage of its structured AnnData output). Give it a try: github.com/COMBINE-lab/...
QCatch: A framework for quality control assessment and analysis of single-cell sequencing data
AbstractMotivation. Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstrea
academic.oup.com
This simulation-cum-benchmark study on MAG making by @tkorem.bsky.social & team looks really interesting. Loads of plots and results to work through! www.biorxiv.org/content/10.6...
biorxiv.org
Accepted to ISMB'26. Revised paper is here: jermp.github.io/assets/pdf/p.... I'd like to thank @robp.bsky.social once again and all the received feedback from the reviewers. To me, ISMB has had the highest quality review process over the past few years!
jermp.github.io
This preprint significantly improves the original design of SSHash to accelerate query and construction time. Looking back at one’s own work is very important and sometimes surprising ☺️
Following up on this - MADRe is now officially published 🎉 Very grateful for the guidance of @msikic.bsky.social @rvicedomini.bsky.social and Kresimir Krizanovic 🔗 academic.oup.com/gigascience/...
I am happy to share our new preprint introducing MADRe - a pipeline for Metagenomic Assembly-Driven Database Reduction, enabling accurate and computationally efficient strain-level metagenomic classification. 🔗https://www.biorxiv.org/content/10.1101/2025.05.12.653324v1 1/9
Myloasm, our long-read metagenome assembler, is now published! w/ @mgmarin.bsky.social and @lh3lh3.bsky.social Very rewarding after > a year of development and countless hours thinking about assembly. Thanks to beta testers, Li lab, and reviewers who gave very helpful feedback. rdcu.be/famFj
High-resolution metagenome assembly for modern long reads with myloasm
Nature Biotechnology - A long-read metagenome assembly method recovers circular and complete genomes better than existing tools.
rdcu.be
High-resolution metagenome assembly for modern long reads with myloasm - @lh3lh3.bsky.social @jimshaw.bsky.social @danafarber.bsky.social @harvardmed.bsky.social go.nature.com/3PBEwvR