A survey of k-mer methods and applications in bioinformatics

Comput Struct Biotechnol J. 2024 May 21:23:2289-2303. doi: 10.1016/j.csbj.2024.05.025. eCollection 2024 Dec.

Abstract

The rapid progression of genomics and proteomics has been driven by the advent of advanced sequencing technologies, large, diverse, and readily available omics datasets, and the evolution of computational data processing capabilities. The vast amount of data generated by these advancements necessitates efficient algorithms to extract meaningful information. K-mers serve as a valuable tool when working with large sequencing datasets, offering several advantages in computational speed and memory efficiency and carrying the potential for intrinsic biological functionality. This review provides an overview of the methods, applications, and significance of k-mers in genomic and proteomic data analyses, as well as the utility of absent sequences, including nullomers and nullpeptides, in disease detection, vaccine development, therapeutics, and forensic science. Therefore, the review highlights the pivotal role of k-mers in addressing current genomic and proteomic problems and underscores their potential for future breakthroughs in research.

Keywords: K-mers; Neomers; Nullomers; Nullpeptides; Primes; Sequence Analysis.

Publication types

  • Review