The impact of different negative training data on regulatory sequence predictions.
Ontology highlight
ABSTRACT: Regulatory regions, like promoters and enhancers, cover an estimated 5-15% of the human genome. Changes to these sequences are thought to underlie much of human phenotypic variation and a substantial proportion of genetic causes of disease. However, our understanding of their functional encoding in DNA is still very limited. Applying machine or deep learning methods can shed light on this encoding and gapped k-mer support vector machines (gkm-SVMs) or convolutional neural networks (CNNs) are commonly trained on putative regulatory sequences. Here, we investigate the impact of negative sequence selection on model performance. By training gkm-SVM and CNN models on open chromatin data and corresponding negative training dataset, both learners and two approaches for negative training data are
SUBMITTER: Krutzfeldt LM
PROVIDER: S-EPMC7707526 | biostudies-literature | 2020
REPOSITORIES: biostudies-literature
ACCESS DATA