Taxonomic classification with maximal exact matches in KATKA kernels and minimizer digests

Draesslerová, Dominika; Ahmed, Omar; Gagie, Travis; Holub, Jan; Langmead, Ben; Manzini, Giovanni; Navarro, Gonzalo

Computer Science > Data Structures and Algorithms

arXiv:2402.06935 (cs)

[Submitted on 10 Feb 2024 (v1), last revised 4 Apr 2024 (this version, v2)]

Title:Taxonomic classification with maximal exact matches in KATKA kernels and minimizer digests

Authors:Dominika Draesslerová, Omar Ahmed, Travis Gagie, Jan Holub, Ben Langmead, Giovanni Manzini, Gonzalo Navarro

View PDF HTML (experimental)

Abstract:For taxonomic classification, we are asked to index the genomes in a phylogenetic tree such that later, given a DNA read, we can quickly choose a small subtree likely to contain the genome from which that read was drawn. Although popular classifiers such as Kraken use $k$-mers, recent research indicates that using maximal exact matches (MEMs) can lead to better classifications. For example, we can build an augmented FM-index over the the genomes in the tree concatenated in left-to-right order; for each MEM in a read, find the interval in the suffix array containing the starting positions of that MEM's occurrences in those genomes; find the minimum and maximum values stored in that interval; take the lowest common ancestor (LCA) of the genomes containing the characters at those positions. This solution is practical, however, only when the total size of the genomes in the tree is fairly small. In this paper we consider applying the same solution to three lossily compressed representations of the genomes' concatenation: a KATKA kernel, which discards characters that are not in the first or last occurrence of any $k_{\max}$-tuple, for a parameter $k_{\max}$; a minimizer digest; a KATKA kernel of a minimizer digest. With a test dataset and these three representations of it, simulated reads and various parameter settings, we checked how many reads' longest MEMs occurred only in the sequences from which those reads were generated ("true positive" reads). For some parameter settings we achieved significant compression while only slightly decreasing the true-positive rate.

Subjects:	Data Structures and Algorithms (cs.DS); Genomics (q-bio.GN); Populations and Evolution (q-bio.PE)
Cite as:	arXiv:2402.06935 [cs.DS]
	(or arXiv:2402.06935v2 [cs.DS] for this version)
	https://doi.org/10.48550/arXiv.2402.06935

Submission history

From: Travis Gagie [view email]
[v1] Sat, 10 Feb 2024 12:20:43 UTC (334 KB)
[v2] Thu, 4 Apr 2024 20:50:46 UTC (434 KB)

Computer Science > Data Structures and Algorithms

Title:Taxonomic classification with maximal exact matches in KATKA kernels and minimizer digests

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Data Structures and Algorithms

Title:Taxonomic classification with maximal exact matches in KATKA kernels and minimizer digests

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators