Gaëtan Benoit

@gaetanbenoit.bsky.social

Postdoc researcher in bioinformatics at Pasteur institute. Scalable methods and software for metagenomics. https://github.com/GaetanBenoitDev

It's out! Excited to present the Great Barrier Reef Microbial Genomes Database (GBR-MGD), a comprehensive DB of 1000s of high-quality prokaryote, virus, plasmid, and chromosome-level eukaryote MAGs using Nanopore long reads. Subthreads incoming. Please share widely. 🙂 www.nature.com/articles/s41...

The planktonic microbiome of the Great Barrier Reef - Nature

The Great Barrier Reef Microbial Genomes Database compiles prokaryotic, viral and eukaryotic genomes from seawater collected from the Great Barrier Reef, providing a rich resource for the study of mar...

nature.com

H

Minibwa is a hybrid of bwa-mem and minimap2 and the successor of bwa-mem for short-read mapping. ~4X/2.5X as fast as bwa-mem/bwa-mem2 for WGS reads at comparable accuracy. Native support of directional bisulfite-seq. Applicable to long reads. Preprint at arxiv.org/abs/2606.15357

Bild

Our method, savont, for generating amplicon sequence variants (ASVs) for long-read amplicons is now on bioRxiv. Work with @lh3lh3.bsky.social and help from @mkddueholm.bsky.social and team (Marie Riisgaard-Jensen, @kirk3gaard.bsky.social, Kasper Skytte Andersen) github.com/bluenote-157... 1/6

bioRxiv Bioinfo@biorxiv-bioinfo.bsky.social · 2mo ago

Sensitive long-read amplicon sequence variant recovery with savont https://www.biorxiv.org/content/10.64898/2026.05.26.727271v1

P

This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.

Hash functions in nucleotide sequence analysis

Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.

doi.org

PPaul Medvedev @pashadag.bsky.social · last yr.

1/4 Hash functions in genomic sequence analysis (tinyurl.com/4kk9ccmt) : a new survey written together with Ke Chen, Xiang Li, Qian Shi, and Mingfu Shao. Before submitting it, we are posting it online to get feedback from the community.

HERRO has been published in @nature.com nature.com/articles/s41.... This achievement is a result of the great work by Dominik Stanojevic, with contributions from Dehui Lin, @sergeynurk.bsky.social, and Paola Florez de Sessions Welcome to the era of high-quality genome assemblies supported by AI.

Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads - Nature

Nature - Telomere-to-Telomere Assembly Using HERRO-Corrected Simplex Nanopore Reads

nature.com

I was investigating the genomes that I didn't manage to convert to near-complete MAGs in my assembly graph (the components in gray). The circle on top left is actually a complete genome but with 40% completeness (both in metaMDBG and myloasm)

Bild