Giulio Ermanno Pibiri

@jermp.bsky.social

Associate Prof. of CS at Ca' Foscari University of Venice. Indexing, Data Compression, Algorithms.

Excited to be at #ISMB2026 this year and to be presenting the work @jermp.bsky.social and I have done on improving the SSHash data structure academic.oup.com/bioinformati.... Presentation is tomorrow, July 13 at 2:20 in the International Ballroom Center. Stop by if you're attending!

Optimizing sparse and skew hashing: faster k-mer dictionaries

AbstractMotivation. Representing a set of k-mers—strings of length k—in small space under fast lookup queries is a fundamental requirement for several appl

academic.oup.com

P

This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.

Hash functions in nucleotide sequence analysis

Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.

doi.org

PPaul Medvedev @pashadag.bsky.social · last yr.

1/4 Hash functions in genomic sequence analysis (tinyurl.com/4kk9ccmt) : a new survey written together with Ke Chen, Xiang Li, Qian Shi, and Mingfu Shao. Before submitting it, we are posting it online to get feedback from the community.

This preprint significantly improves the original design of SSHash to accelerate query and construction time. Looking back at one’s own work is very important and sometimes surprising ☺️

Rob Patro@robp.bsky.social · 7mo ago

Very excited about this latest work led by @jermp.bsky.social! Since it's initial release, SSHash has served as the basis for several other tools (Fulgor, piscem, etc.). It was already very fast. It is now *substantially* faster! www.biorxiv.org/content/10.6...

Great work and congrats to all authors! 🥳 Can't wait to read the preprint. Just for reference: a Fulgor index ⚡ on the same HPRC collection takes 8.26 GB (not in its most succinct representation). Fulgor, however, can be regarded as a "lossy" method here since it is based on kmers.

Mohsen Zakeri@mohsenzakeri.bsky.social · 10mo ago

5/6 On the 466 haplotypes from the 2nd release of HPRC, the fastest Movi 2 index is under 50 GB. It can be reduced to 24 GB while remaining over 3x faster than SPUMONI. Movi 2 is smaller and faster than ropebwt3, although it computes PMLs, which are easier to get than the SMEMs found by ropebwt3.

Hi bioinformatics, genomics and CS friends! Please help me spread the word. I'm hiring a postdoc! Come work on cutting edge method development in algorithmic genomics with me and my group at @umdscience.bsky.social! 🖥️🧬

Rob Patro@robp.bsky.social · 10mo ago

And it's posted! If you're interested and eligible, please consider applying through the UMD portal: umd.wd1.myworkdayjobs.com/en-US/UMCP/j.... If you're a PI working in algorithmic genomics (& you can recommend my lab to your top graduating students ;P), please let them know!