1/ 🧬 & 🖥️ New tool: kmhelpers — a Python toolkit dedicated to kmindex, automating the building, updating & querying of large-scale k-mer genomic indexes. Turning raw sequencing samples into a fast, size-optimized search engine — without needing to be a Bloom-filter expert. 🧵
Pierre Peterlongo
@pierrepeterlongo.bsky.social
Inria Senior researcher. Head of the https://team.inria.fr/genscale/ at Inria and Irisa. Algorithmics for sequencing data analyses, genomics and metagenomics.
Might be useful. Here is a short blog entry describing a tool that annotates a PDF using a text file. I couldn't find anything that simply does this single task, so I created this. Blog post: pierrepeterlongo.github.io Git repo: github.com/pierrepeterl...
minibwa v0.3 released with a few minor bug fixes and two missing bwa-mem features (XA tag for secondary hits and option -H to inject header lines). Also added the "mem" subcommand to mimic "bwa mem" CLI to some extent. github.com/lh3/minibwa/...
Release Minibwa-0.3 (r391) · lh3/minibwa
Notable changes: New feature: added the mem subcommand to mimic the bwa-mem command-line interface (CLI). Most input/output options and commonly used options are retained; unsupported or incompat...
github.com
Movi 2 has appeared (as an advance article) in Bioinformatics 🧬 Faster, leaner pangenome queries — half the memory of Movi 1, ~30% faster. Paper: academic.oup.com/bioinformati... Code: github.com/mohsenzakeri/Movi (1/6)
Validate User
academic.oup.com
Happy to see that K2Rmini was recommended today by PCI Mathematical & Computational Biology. "quickly evaluate whether an arbitrary sequence has a number of k-mer [of interest] matches above or below a threshold." by @imartayan.bsky.social and colleagues: www.biorxiv.org/content/10.1...
📄 OrthoFinder: improved phylogenetic orthology inference with enhanced accuracy and scalability 🧬🖥️ ⬇️ www.nature.com/articles/s41...
OrthoFinder: improved phylogenetic orthology inference with enhanced accuracy and scalability - Nature Methods
The updated OrthoFinder v3 software boosts accuracy and scalability in phylogenetic orthology inference with massive and diverse datasets.
nature.com
🌎 🧬 🖥️ logan-search.org the tool to query all SRA sequences (Dec 2023 snapshot) — just got major updates. Here's what's new. 🧵 1/12
1/8 🛠️ Introducing logan_blaster — a command-line tool to locally align your sequences against Logan contigs or unitigs. github.com/pierrepeterlongo/logan_blaster
GitHub - pierrepeterlongo/logan_blaster: Given either a kmviz id (obtained from a Logan-Search query) or a query (fasta sequence), and a file containing a list of SRA accessions (provided or not by Lo...
Given either a kmviz id (obtained from a Logan-Search query) or a query (fasta sequence), and a file containing a list of SRA accessions (provided or not by Logan-Search results) run a local blast ...
github.com
🔍 New paper in Bioinformatics Advances: "Kaminari: A frugal colored index for approximate k-mer queries" Read it here: https://doi.org/10.1093/bioadv/vbag120 Authors include: @yhhshb.bsky.social, @yoann.bsky.social, @robp.bsky.social, @pierrepeterlongo.bsky.social, @jermp.bsky.social
Good Friday Evening news: we updated back_to_sequences (find the origin of kmers) - Faster - Can consider multiline fasta files - Much easier installation: see github.com/pierrepeterl...
The Metagraph paper is out in Nature; it showed up in my feeds today! Congratulations to Mikhail Karasikov, @gxxxr.bsky.social, @akkah21.bsky.social and all of the other authors (whom I'd love to follow on Bluesky if I can find you ;P) www.nature.com/articles/s41...
Efficient and accurate search in petabase-scale sequence repositories - Nature
MetaGraph enables scalable indexing of large sets of DNA, RNA or protein sequences using annotated de Bruijn graphs.
nature.com
Preprint out for myloasm, our new nanopore / HiFi metagenome assembler! Nanopore's getting accurate, but 1. Can this lead to better metagenome assemblies? 2. How, algorithmically, to leverage them? with co-author Max Marin @mgmarin.bsky.social, supervised by Heng Li @lh3lh3.bsky.social 1 / N
High-resolution metagenome assembly for modern long reads with myloasm https://www.biorxiv.org/content/10.1101/2025.09.05.674543v1
❗ I clearly consider this result as THE most important result achieved over this last decade for exploiting and democratizing genomic data. I think there will be a "before" and an "after" logan and logan-search github.com/IndexThePlan... logan-search.org Have a look at this thread
🌎👩🔬 For 15+ years biology has accumulated petabytes (million gigabytes) of🧬DNA sequencing data🧬 from the far reaches of our planet.🦠🍄🌵 Logan now democratizes efficient access to the world’s most comprehensive genetics dataset. Free and open. doi.org/10.1101/2024...
📜 Excited to share insights from our recent paper: "Kaminari: a resource-frugal index for approximate colored k-mer queries". The study aims to efficiently identify documents containing a query string, focusing on DNA strings. www.biorxiv.org/content/10.1... 🧬 🖥️ 1/8
Maybe the simplest idea to decrease overestimations of a counting bloom filter. A trivial observation + 10 lines of code. I'm surprised it has not been described before. Please comment if this is not the case. Blog post here: pierrepeterlongo.github.io/2025/03/17/m... 🧪🧬🖥️
Today I wanted to know the number of unique 27-mers in the hg38 human genome (spoiler there are 2.49 billion). I found no tool for doing this. So I wrote that github.com/pierrepeterl... It may help. Please use it / improve it. 🧬💻 #bioinformatics
GitHub - pierrepeterlongo/unique_kmer_counter: Count number of unique kmers from fasta or fasta.gz files
Count number of unique kmers from fasta or fasta.gz files - pierrepeterlongo/unique_kmer_counter
github.com
We are back in the Town Theatre for a great lecture on Alignment, by @rayanchikhi.bsky.social! 🧬💻 #evomics2025 #genomics #bioinformatics
bsky.app/profile/pier... Applications for this position are still open. If you're passionate about large-scale science, we'd love to hear from you. 🧬 & 🖥️
🚨🚨🚨 We are hiring 🚨🚨🚨 After the creation of logan-search (see: bsky.app/profile/pier...) we propose a 2-years engineer position for continuing the development and optimizations. With @rayanchikhi.bsky.social and @tlemane.bsky.social Details + applications: recrutement.inria.fr/public/class...
🚨🚨🚨 We are hiring 🚨🚨🚨 After the creation of logan-search (see: bsky.app/profile/pier...) we propose a 2-years engineer position for continuing the development and optimizations. With @rayanchikhi.bsky.social and @tlemane.bsky.social Details + applications: recrutement.inria.fr/public/class...
🧬🔍There are 50 petabases of freely-available DNA sequencing data. We introducing Logan Search which allows you to search for any DNA sequence in minutes, bringing Earth’s largest genomic resource to your fingertips. 🏔️ logan-search.org 🏔️ #Genomics #Bioinformatics #OpenScience
🚨 Call for Papers: RECOMB-seq 2025 🚨 🗓️ Dates: April 24-25, 2025 📍 Location: Seoul, South Korea Key deadlines: 🔹 Abstract registration: Jan 24, 2025 🔹 Submission: Jan 31, 2025 More details: recomb-seq.github.io/papers/
Call for Papers
RECOMB-seq is the RECOMB Satellite Conference on Biological Sequence Analysis
recomb-seq.github.io
🗓️Tomorrow, Friday December 13, Khodor HANNOUSH, from the @genscaleteam.bsky.social team, will defend his thesis entitled “Dynamic Pan-genome Graphs”. Details by following this link: www.irisa.fr/date/2024-12...
Graphes dynamiques de pangénome | le site web de l'IRISA
irisa.fr
Amazing ideas here www.biorxiv.org/content/bior... from @yoann.bsky.social and collaborators. Reorganize minimizers to allow kmers dichotomic search. That's brilliant. #bioinformatics 🧬🖥️
I made a starter pack for algorithmic genomics. It's certainly incomplete, but already has a ton of awesome peeps. Let me know if you know people I should add (with a focus on algorithms and data structures in genomics) go.bsky.app/TRWCnZs
🧬🔍There are 50 petabases of freely-available DNA sequencing data. We introducing Logan Search which allows you to search for any DNA sequence in minutes, bringing Earth’s largest genomic resource to your fingertips. 🏔️ logan-search.org 🏔️ #Genomics #Bioinformatics #OpenScience