Unknown

Dataset Information

0

Scalable neighbour search and alignment with uvaia.


ABSTRACT: Despite millions of SARS-CoV-2 genomes being sequenced and shared globally, manipulating such data sets is still challenging, especially selecting sequences for focused phylogenetic analysis. We present a novel method, uvaia, which is based on partial and exact sequence similarity for quickly extracting database sequences similar to query sequences of interest. Many SARS-CoV-2 phylogenetic analyses rely on very low numbers of ambiguous sites as a measure of quality since ambiguous sites do not contribute to single nucleotide polymorphism (SNP) differences. Uvaia overcomes this limitation by using measures of sequence similarity which consider partially ambiguous sites, allowing for more ambiguous sequences to be included in the analysis if needed. Such fine-grained definition of similarity allows not only for better phylogenetic analyses, but could also lead to improved classification and biogeographical inferences. Uvaia works natively with compressed files, can use multiple cores and efficiently utilises memory, being able to analyse large data sets on a standard desktop.

SUBMITTER: de Oliveira Martins L 

PROVIDER: S-EPMC10924453 | biostudies-literature | 2024

REPOSITORIES: biostudies-literature

altmetric image

Publications

Scalable neighbour search and alignment with uvaia.

de Oliveira Martins Leonardo L   Mather Alison E AE   Page Andrew J AJ  

PeerJ 20240306


Despite millions of SARS-CoV-2 genomes being sequenced and shared globally, manipulating such data sets is still challenging, especially selecting sequences for focused phylogenetic analysis. We present a novel method, uvaia, which is based on partial and exact sequence similarity for quickly extracting database sequences similar to query sequences of interest. Many SARS-CoV-2 phylogenetic analyses rely on very low numbers of ambiguous sites as a measure of quality since ambiguous sites do not c  ...[more]

Similar Datasets

| S-EPMC10491953 | biostudies-literature
| S-EPMC4963998 | biostudies-literature
| S-EPMC3311098 | biostudies-literature
| S-EPMC8523058 | biostudies-literature
| S-EPMC3228549 | biostudies-literature
| S-EPMC2951093 | biostudies-literature
| S-EPMC6030823 | biostudies-literature
| S-EPMC4271471 | biostudies-literature
| S-EPMC2859133 | biostudies-literature
| S-EPMC4986259 | biostudies-literature