Unknown

Dataset Information

0

RNAVirHost: a machine learning-based method for predicting hosts of RNA viruses through viral genomes.


ABSTRACT: The high-throughput sequencing technologies have revolutionized the identification of novel RNA viruses. Given that viruses are infectious agents, identifying hosts of these new viruses carries significant implications for public health and provides valuable insights into the dynamics of the microbiome. However, determining the hosts of these newly discovered viruses is not always straightforward, especially in the case of viruses detected in environmental samples. Even for host-associated samples, it is not always correct to assign the sample origin as the host of the identified viruses. The process of assigning hosts to RNA viruses remains challenging due to their high mutation rates and vast diversity. In this study, we introduce RNAVirHost, a machine learning-based tool that predicts the hosts of RNA viruses solely based on viral genomes. RNAVirHost is a hierarchical classification framework that predicts hosts at different taxonomic levels. We demonstrate the superior accuracy of RNAVirHost in predicting hosts of RNA viruses through comprehensive comparisons with various state-of-the-art techniques. When applying to viruses from novel genera, RNAVirHost achieved the highest accuracy of 84.3%, outperforming the alignment-based strategy by 12.1%. The application of machine learning models has proven beneficial in predicting hosts of RNA viruses. By integrating genomic traits and sequence homologies, RNAVirHost provides a cost-effective and efficient strategy for host prediction. We believe that RNAVirHost can greatly assist in RNA virus analyses and contribute to pandemic surveillance.

SUBMITTER: Chen G 

PROVIDER: S-EPMC11340644 | biostudies-literature | 2024 Jan

REPOSITORIES: biostudies-literature

altmetric image

Publications

RNAVirHost: a machine learning-based method for predicting hosts of RNA viruses through viral genomes.

Chen Guowei G   Jiang Jingzhe J   Sun Yanni Y  

GigaScience 20240101


<h4>Background</h4>The high-throughput sequencing technologies have revolutionized the identification of novel RNA viruses. Given that viruses are infectious agents, identifying hosts of these new viruses carries significant implications for public health and provides valuable insights into the dynamics of the microbiome. However, determining the hosts of these newly discovered viruses is not always straightforward, especially in the case of viruses detected in environmental samples. Even for ho  ...[more]

Similar Datasets

| S-EPMC3235098 | biostudies-literature
| S-EPMC7086167 | biostudies-literature
| S-EPMC8611875 | biostudies-literature
| S-EPMC6536379 | biostudies-literature
2021-07-09 | GSE163896 | GEO
| S-EPMC10805179 | biostudies-literature
| S-EPMC8087038 | biostudies-literature
| S-EPMC11343755 | biostudies-literature
| S-EPMC10691042 | biostudies-literature
| S-EPMC10319836 | biostudies-literature