Ontology highlight
ABSTRACT: Summary
Bioinformatics applications increasingly rely on ad hoc disk storage of k-mer sets, e.g. for de Bruijn graphs or alignment indexes. Here, we introduce the K-mer File Format as a general lossless framework for storing and manipulating k-mer sets, realizing space savings of 3-5× compared to other formats, and bringing interoperability across tools.Availability and implementation
Format specification, C++/Rust API, tools: https://github.com/Kmer-File-Format/.Supplementary information
Supplementary data are available at Bioinformatics online.
SUBMITTER: Dufresne Y
PROVIDER: S-EPMC9477520 | biostudies-literature | 2022 Sep
REPOSITORIES: biostudies-literature

Dufresne Yoann Y Lemane Teo T Marijon Pierre P Peterlongo Pierre P Rahman Amatur A Kokot Marek M Medvedev Paul P Deorowicz Sebastian S Chikhi Rayan R
Bioinformatics (Oxford, England) 20220901 18
<h4>Summary</h4>Bioinformatics applications increasingly rely on ad hoc disk storage of k-mer sets, e.g. for de Bruijn graphs or alignment indexes. Here, we introduce the K-mer File Format as a general lossless framework for storing and manipulating k-mer sets, realizing space savings of 3-5× compared to other formats, and bringing interoperability across tools.<h4>Availability and implementation</h4>Format specification, C++/Rust API, tools: https://github.com/Kmer-File-Format/.<h4>Supplementar ...[more]