Benjamin J. Buchfink

@bbuchfink.bsky.social

Independent scientist, Tübingen, Germany. Developer of the DIAMOND protein aligner. https://github.com/bbuchfink/diamond https://www.linkedin.com/in/benjamin-j-buchfink-875692105

Bioinformatics rewrites miss the relational impact tofolks that are still actively maintaining and developing the software. I've been guilty of this myself, and I'll be sharing my story soon so others can learn from it.

A few additional thoughts following my post: 1. There's likely far more biological data than we often assume. The challenge's that it remains highly fragmented. Large international efforts to integrate them could greatly advance understanding of complex, dynamic biological systems.

César de la Fuente@delafuentelab.bsky.social · 2w ago

Biology has plenty of data—the challenge is making it usable. AllTheBacteria transforms 2.44 million public bacterial and archaeal genomes into an open, uniformly processed, searchable, AI-ready resource. www.biorxiv.org/content/10.1...

This is awful to hear, describing how Sean Eddy (HMMER, infernal, pfam, rfam) has been defunded. The letter said his work "had been determined to be of absolutely no value to the US taxpayer, and therefore it was being specifically terminated," www.npr.org/2026/05/21/n...

Researchers say the Trump administration is finding new ways to punish science

Even with federal grants largely restored, scientists say the Trump administration is still preventing those funds from reaching them. The consequences, they say, are already becoming clear.

npr.org

P

This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.

Hash functions in nucleotide sequence analysis

Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.

doi.org

PPaul Medvedev @pashadag.bsky.social · last yr.

1/4 Hash functions in genomic sequence analysis (tinyurl.com/4kk9ccmt) : a new survey written together with Ke Chen, Xiang Li, Qian Shi, and Mingfu Shao. Before submitting it, we are posting it online to get feedback from the community.

I'm not looking forward to a future where all the tools are being vibe-rewritten into languages people don't want to learn. Who will maintain all this? The original maintainers won't. It's not the language they were comfortable with. Does the prompter understand the tool well enough?

Phil Ewels@ewels.bsky.social · 4mo ago

Super excited to be launching two things today: #RustQC 🦀🧬 and rewrites.bio 🚀 I used AI to rewrite 15 RNA-seq QC tools into a single Rust binary (I've never written any Rust). It ended up being over 60x faster. Here's the story 🧵 seqeralabs.github.io/RustQC/

The 2026 Workshop on Genomics comes to an end! It has been two intense and inspiring weeks of Genomics in Český Krumlov. We hope everyone is going back home with a renewed excitement for science and new friends and collaborations around the world! 🙆🏻‍♀️🧬 #evomics2026

Bild