Unknown

Dataset Information

Comprehensive annotations of the mutational spectra of SARS-CoV-2 spike protein: a fast and accurate pipeline.


ABSTRACT: Infecting millions of people, the SARS-CoV-2 is evolving at an unprecedented rate, demanding advanced and specified analytic pipeline to capture the mutational spectra. In order to explore mutations and deletions in the spike (S) protein - the most-discussed protein of SARS-CoV-2 - we comprehensively analyzed 35,750 complete S protein-coding sequences through a custom Python-based pipeline. This GISAID-collected dataset of until 24 June 2020 covered six continents and five major climate zones. We identified 27,801 (77.77% sequences) mutated strains compared to reference Wuhan-Hu-1 wherein 84.40% of these strains mutated by only a single amino acid (aa). An outlier strain (EPI_ISL_463893) from Bosnia and Herzegovina possessed six aa substitutions. We also identified 11 residues with high aa

SUBMITTER: Rahman MS 

PROVIDER: S-EPMC7646266 | biostudies-literature | 2021 May

REPOSITORIES: biostudies-literature

altmetric image

Publications

Sorry, this publication's infomation has not been loaded in the Indexer, please go directly to PUBMED or Altmetric.

Similar Datasets