Athena: Automated Tuning of k-mer based Genomic Error Correction Algorithms using Language Models.
Ontology highlight
ABSTRACT: The performance of most error-correction (EC) algorithms that operate on genomics reads is dependent on the proper choice of its configuration parameters, such as the value of k in k-mer based techniques. In this work, we target the problem of finding the best values of these configuration parameters to optimize error correction and consequently improve genome assembly. We perform this in an adaptive manner, adapted to different datasets and to EC tools, due to the observation that different configuration parameters are optimal for different datasets, i.e., from different platforms and species, and vary with the EC algorithm being applied. We use language modeling techniques from the Natural Language Processing (NLP) domain in our algorithmic suite, Athena, to automatically tune the perfor
SUBMITTER: Abdallah M
PROVIDER: S-EPMC6834855 | biostudies-literature | 2019 Nov
REPOSITORIES: biostudies-literature
ACCESS DATA