Accelerating k-mer-based sequence filtering

Martayan I., Vandamme L., Constantinides B., Cazaux B., Paperman C., Limasset A.

Motivation. The exponential growth of global sequencing data repositories presents both analytical challenges and opportunities. While k-mer-based indexing has improved scalability over traditional alignment for identifying relevant documents, pinpointing the exact sequences matching numerous queries remains a hurdle. In particular, search-ing for numerous k-mers with a single large query or multiple distinct queries strains existing exact matching tools, whose performance scales poorly with an increasing number of patterns. At the same time, indexing entire vast datasets for infrequent or ad-hoc searches is often resource-prohibitive. Designing fast methods for matching a large number of k-mers without exhaustive pre-indexing is therefore critical. Contributions. We propose an efficient solution to the problem of k-mer-based sequence filter-ing: given a set of k-mers of interests and a threshold, quickly evaluate whether an ar-bitrary sequence has a number of k-mer matches above or below the threshold. Our approach demonstrates how minimizer-based based sketching, alongside SIMD accel-eration, can enhance the performance of streaming searches, and is implemented as a Rust tool named K2Rmini. On a consumer laptop, K2Rmini is able to filter long reads at 2 Gbp/s.

DOI

10.24072/pcjournal.735

Type

Journal article

Publication Date

2026-01-01T00:00:00+00:00

Volume

6

Permalink More information Close