This will make indexing k-mers in e.g. HPRC (human) and GTDB (bacteria) a breeze
I've been working on Pandedup the past weeks: a tool to quickly dedup the k-mer content of a pangenome (read: HPRCv2). Just uploaded a k=64 SPSS (k-mer spectrum preserving string set) as well: zenodo.org/records/2172... This reduces the ~1.5Tbp input to just 10.6Gbp. github.com/RagnarGrootK...