For the past 30 years, โwhole-genome sequencingโ has been a misnomer. Today the T2T Consortium publishes a dozen papers heralding a future of truly complete genomes for humans and nearly any vertebrate ๐จโ๐ฌ๐๐ฆ๐๐ฆ๐๐ซ๐น๐ (sorry, no salamanders): www.cell.com/consortium/t... ๐งต[1/15]
Can Firtina
@firtinac.bsky.social
Assistant Professor of Computer Science at UMD | PhD from ETH Zurich
I got frustrated that using piscem-rs to map 10x Flex reads was entirely dominated by decompressing the single pair of input files. So, GPT5.6-Sol & I built a thing! Rapidgzip-rust (rapid-gzip algo in rust) github.com/COMBINE-lab/... 2.3 billion flex reads in <2 min? Parsing w paraseq. Yes plz!
๐ข Genome Informatics 2026 (Hinxton UK + virtual, 2โ4 Dec) is coming! Some confirmed speakers now listed: coursesandconferences.wellcomeconnectingscience.org/event/genome... Early-bird registration & bursary deadlines: 7 Sept. Abstract deadline: 5 Oct. Please submit your work & join us!
Very happy to be presenting "Taxonomic classification using Kraken 2" in today's #ISMB2026 tutorial "Metagenomic sequence analysis using k-mer based methods": www.iscb.org/ismb2026/wha... In about 15 minutes!
Tutorials - ISMB 2026
ISCB - International Society for Computational Biology
iscb.org
Some professional news: After adding use MoreResearch as _; use MoreMentoring as _; use MoreTeaching as _; use MoreService as _; the UMD trait solver has accepted the where-clause where Rob: Professor On this occasion, some thanks and some thoughts are in order! 1/5
salmon 1.12.0 is out (and on bioconda)! There are some important bug fixes and improvements. Also, this is the last planned C++ release of Salmon! github.com/COMBINE-lab/... "The last? What's happening, Rob?" 1/2
Release salmon 1.12.0 โ final C++ release ยท COMBINE-lab/salmon
salmon 1.12.0 โ release notes This is the final release of the C++ implementation of salmon. Future development continues in salmon 2.0, a from-scratch Rust rewrite that is faster, easier to build ...
github.com
๐ ๐งฌ ๐ฅ๏ธ logan-search.org the tool to query all SRA sequences (Dec 2023 snapshot) โ just got major updates. Here's what's new. ๐งต 1/12
Memory-safe high-performance sequence mapping with rammap https://www.biorxiv.org/content/10.64898/2026.05.26.726289v1
Revamped the teaching materials site. Light/dark mode switch at the top. More consistent, compact presentation of all the materials. + Icons and such! www.langmead-lab.org/teaching.html
๐ข ๐ช๐๐๐ ๐๐๐ ๐ท๐๐๐๐๐๐๐๐๐๐๐๐: ๐ต๐๐๐-๐ฎ๐๐๐๐๐๐๐๐๐ ๐จ๐ ๐๐๐๐๐๐๐ ๐ช๐๐๐๐๐๐๐๐ ๐๐๐ ๐ถ๐๐๐๐ @ ๐ญ๐ช๐ช๐ด ๐๐๐๐ ๐ Workshop Website: events.safari.ethz.ch/fccm2026-omi... ๐๏ธ ๐๐ก๐๐ง ๐๐ง๐ ๐๐ก๐๐ซ๐: 16th May 2026 (Saturday), Atlanta, Georgia, USA ๐ฅ Livestream: www.youtube.com/live/jUJqOY7... @omutlu.bsky.social @firtinac.bsky.social
Next-Generation Adaptable Computing for Omics @ FCCM 2026
FCCM 2026: 2nd Workshop on Next-Generation Adaptable Computing for Omics https://events.safari.ethz.ch/fccm2026-omics/ Date and Time: May 16, from 9:00 AM (EDT) to 01:00 PM (EDT) Organizers: Prof. Ma...
youtube.com
This is now published in Genome Research (doi.org/10.1101/gr.2...). Thank you everyone for your feedback and also the anonymous reviewers who helped to greatly improve the paper. I hope this becomes a useful resource for the community.
Hash functions in nucleotide sequence analysis
Randomness is a powerful tool in the design and analysis of algorithms and data structures for nucleotide sequence data. Nucleotide sequences are not themselves random but are often randomized using hash functions. Despite their widespread use in genomics, there is no comprehensive review of the types of hash functions used and their various applications. In this survey intended for bioinformatic methods developers, we divide hash functions into four categories: scattering hash functions, permutations, minimum perfect hash functions, and locality-sensitive hash functions. For each category, we provide examples of both general-use hash functions that have been applied in nucleotide sequence analysis and hash functions that have been designed specifically for nucleotide sequence analysis. We highlight their salient properties, commonalities, differences, and application areas.
doi.org
1/4 Hash functions in genomic sequence analysis (tinyurl.com/4kk9ccmt) : a new survey written together with Ke Chen, Xiang Li, Qian Shi, and Mingfu Shao. Before submitting it, we are posting it online to get feedback from the community.
The RECOMB-Seq 2026 program is now available! Join us May 24โ25 in Thessaloniki, Greece, for two days of cutting-edge biological sequence analysis, with keynotes by Camille Marchet (CNRS) and Manolis Kellis (MIT). Full schedule: recomb-seq.github.io/seq2026/prog... #RECOMBseq
Program
RECOMB-Seq 2026 Web Page
recomb-seq.github.io
QCatch is now published in Bioinformatics (academic.oup.com/bioinformati...)! Great work from Yuan and Dongze for quality control and analysis downstream of simpleaf/alevin-fry (taking advantage of its structured AnnData output). Give it a try: github.com/COMBINE-lab/...
QCatch: A framework for quality control assessment and analysis of single-cell sequencing data
AbstractMotivation. Single-cell sequencing data analysis requires robust quality control (QC) to mitigate technical artifacts and ensure reliable downstrea
academic.oup.com
๐ Congratulations to Onur Mutlu on being named 2025 AAAS Fellow "for foundational and innovative contributions to computer engineering research, education, and practice, especially in computer architecture and memory and storage systemsโ safari.ethz.ch/onur-mutlu-n... @omutlu.bsky.social @aaas.org
Simon Ambrozak, Ulysse McConnell, Bhargav Srinivasan, Burak Ozkan, Can Firtina: CERN: Correcting Errors in Raw Nanopore Signals Using Hidden Markov Models https://arxiv.org/abs/2603.20420 https://arxiv.org/pdf/2603.20420 https://arxiv.org/html/2603.20420
๐ฃ๐ผ๐๐๐ฑ๐ผ๐ฐ ๐ฎ๐ป๐ฑ ๐ฃ๐ต๐ ๐ฝ๐ผ๐๐ถ๐๐ถ๐ผ๐ป๐ ๐ถ๐ป ๐๐ผ๐บ๐ฝ๐๐๐ฎ๐๐ถ๐ผ๐ป๐ฎ๐น ๐๐ฒ๐ป๐ผ๐บ๐ถ๐ฐ๐ / ๐๐น๐ด๐ผ๐ฟ๐ถ๐๐ต๐บ๐ถ๐ฐ ๐๐ถ๐ผ๐ถ๐ป๐ณ๐ผ๐ฟ๐บ๐ฎ๐๐ถ๐ฐ๐ I am currently recruiting for both: ๐น Postdoc position su.varbi.com/what:job/job... ๐น PhD position su.varbi.com/en/what:job/... Please share with anyone who might be interested!
Postdoktor i Berรคkningsbiologi
Matematiska institutionen bestรฅr av cirka 120 forskare, lรคrare och administrativ personal och รคr organiserad i tre huvudsakliga avdelningar: Matematik, Matematisk statistik och Berรคkningsmatematik
su.varbi.com
This is an important step in the right direction. The methods are improving to start collectively thinking about how we should analyze raw nanopore signals for various new applications that are not possible without signal analysis. Congratulations to the authors!
Transformer-based AI has boosted @nanoporetech.com sequencing accuracy, but at a cost to portability due to GPU demands. Our new work, spearheaded by Sara Bakic, introduces Campolina link.springer.com/article/10.1... to improve nanopore signal segmentation for event-based mappers.
Transformer-based AI has boosted @nanoporetech.com sequencing accuracy, but at a cost to portability due to GPU demands. Our new work, spearheaded by Sara Bakic, introduces Campolina link.springer.com/article/10.1... to improve nanopore signal segmentation for event-based mappers.
Campolina: a deep neural framework for accurate segmentation of nanopore signals - Genome Biology
Nanopore sequencing enables real-time, long-read analysis by processing raw signals as they are produced. A key step, segmentation of signals into events, is typically handled algorithmically, struggl...
link.springer.com
1/ Our paper on Multi-Context Seeds is now out, with @tolyan.bsky.social spearheading the work and contributions from Nicolas and @marcelm.net. We introduce a new seeding concept that improves read alignment accuracy while maintaining speed. link.springer.com/article/10.1...
Multi-context seeds enable fast and high-accuracy read mapping - Genome Biology
A key step in sequence similarity search is to identify shared seeds between a query and a reference sequence. A well-known tradeoff is that longer seeds offer fast searches but reduce sensitivity in ...
link.springer.com
This looks really interesting and perhaps might enable faster overlapping for HERRO @msikic.bsky.social , not sure academic.oup.com/bioinformati... , also not sure Rawsamble can overlap repetitive sequences, for which minimap2 can, but seems to slow down a lot compared to default settings
Rawsamble: Overlapping Raw Nanopore Signals using a Hash-based Seeding Mechanism
AbstractMotivation. Raw nanopore signal analysis is a common approach in genomics to provide fast and resource-efficient analysis without translating the s
academic.oup.com
Iโve written a post about my recent experiences (successes) with AI coding models; the experiences that caused me to re-evaluate my initial judgements, the surprise I had at what can be accomplished, & some fears I have about these tools. Discussion welcome! combine-lab.github.io/blog/2026/02...
COMBINE-lab - The skepticโs guide to generative AI assisted coding
An easy-to-use, flexible website template for labs, with automatic citations, GitHub tag imports, pre-built components, and more.
combine-lab.github.io
๐งฌ Hiring Postdocs at @astar_gis ! We need Computer Scientists and Computational Biologists to develop novel algorithms for de novo assembly of cancer genomes or to help us reconstruct them. Experience in sequence alignment/assembly algorithms or assembly of complex genomes required Please RT! ๐
Happy to share I recently received the ETH Doctoral Medal. More information on LinkedIn: www.linkedin.com/posts/canfir... @safari-eth.bsky.social @omutlu.bsky.social
I am humbled to receive the ETH Doctoral Medal in 2025. This award is given in recognition of the best doctoral theses at ETH Zurich. I am deeply grateful to my PhD advisor, Prof. Onur Mutlu, forโฆ | ...
I am humbled to receive the ETH Doctoral Medal in 2025. This award is given in recognition of the best doctoral theses at ETH Zurich. I am deeply grateful to my PhD advisor, Prof. Onur Mutlu, for his...
linkedin.com
Iโm recruiting a postdoc to work on algorithms for cancer genome reconstruction. We have access to a rich set of tumour samples sequenced across multiple technologies. If interested, feel free to DM. Please share.
@wytamma.bsky.social : so, it took a little bit of extra time (not the flight back from the CZI meeting), but I decided to just f#&$ing do it, and the basic code to build and parse with the auxiliary fastq index is working (github.com/COMBINE-lab/...). 1/2
GitHub - COMBINE-lab/mim: A small, auxiliary index to massively improve parallel fastq parsing
A small, auxiliary index to massively improve parallel fastq parsing - COMBINE-lab/mim
github.com
Our department (Comp Sci) at UMD is hiring this cycle. We have an open-rank search (umd.wd1.myworkdayjobs.com/en-US/UMCP/j...). Consider applying to join our department; weโre a pretty cool group if I say so myself โบ๏ธ!
Assistant Professor, Associate Professor, Professor
Job Description Summary Organization's Summary Statement: The Department of Computer Science is top-ranked for research and teaching, with its undergraduate computer science program ranked 9th among p...
umd.wd1.myworkdayjobs.com
Have you recently completed (or finishing soon) a PhD in CS or a related discipline? Do you want to do research advancing the theory & practice of algorithmic genomics & build tools that people love to use? I'll be looking to hire a postdoc! Official ad coming soon: docs.google.com/document/d/1...
Postdoc Description.docx
Title: Postdoctoral Associate Summary statement: The postdoctoral research associate is responsible for developing novel computational methodology for high-throughput sequence genomics tasks, as well ...
docs.google.com
RawBench: A Comprehensive Benchmarking Framework for Raw Nanopore Signal Analysis Techniques https://www.biorxiv.org/content/10.1101/2025.10.04.680405v1
Thank you folks for your feedback on our survey about Hash functions in genomic sequence analysis. We've updated the paper and you can see the new version here: tinyurl.com/4kk9ccmt.
Dropbox
tinyurl.com
1/4 Hash functions in genomic sequence analysis (tinyurl.com/4kk9ccmt) : a new survey written together with Ke Chen, Xiang Li, Qian Shi, and Mingfu Shao. Before submitting it, we are posting it online to get feedback from the community.
๐๐ฉโ๐ฌ For 15+ years biology has accumulated petabytes (million gigabytes) of๐งฌDNA sequencing data๐งฌ from the far reaches of our planet.๐ฆ ๐๐ต Logan now democratizes efficient access to the worldโs most comprehensive genetics dataset. Free and open. doi.org/10.1101/2024...