Unknown

Dataset Information

0

SCOPE++: sequence classification of homoPolymer emissions.


ABSTRACT: mRNA polyadenylation, the addition of a poly(A) tail to the 3'-end of pre-mRNA, is a process critical to gene expression and regulation in eukaryotes. To understand the molecular mechanisms governing polyadenylation and other relevant biological processes, it is important to identify these poly(A) tails accurately in transcriptome sequencing data and differentiate them from artificial adapter sequences added in the sequencing process. But the annotation of these tails is complicated by the presence of sequencing errors and post-transcriptional modifications. While determining that a tail is present in a given transcript fragment is straight-forward, these obfuscations make the problem of boundary identification a challenge; conventional seed-and-extend algorithms struggle to accurately identify these poly(A) tail end-points. Further, all existing tools that we are aware of focus exclusively on the trimming of poly(A) tails, failing to provide the detailed information needed for studying the polyadenylation process.We have created SCOPE++, an open-source tool for finding the precise border of poly(A) tails and other homopolymers in raw mRNA sequence reads. Based on a Hidden Markov Model (HMM) approach, SCOPE++ accurately identifies specific homopolymer sequences in error-prone EST/cDNA data or RNA-Seq data at a speed appropriate for large sequence sets.We demonstrate that our tool can precisely identify poly(A) tails with near perfect accuracy at the speed required for high-throughput applications, providing a valuable resource for polyadenylation research.

SUBMITTER: Morton JT 

PROVIDER: S-EPMC4165746 | biostudies-literature | 2014 Sep

REPOSITORIES: biostudies-literature

altmetric image

Publications

SCOPE++: sequence classification of homoPolymer emissions.

Morton James T JT   Abrudan Patricia P   Figueroa Nathanial N   Liang Chun C   Karro John E JE  

Genomics 20140801 3


<h4>Background</h4>mRNA polyadenylation, the addition of a poly(A) tail to the 3'-end of pre-mRNA, is a process critical to gene expression and regulation in eukaryotes. To understand the molecular mechanisms governing polyadenylation and other relevant biological processes, it is important to identify these poly(A) tails accurately in transcriptome sequencing data and differentiate them from artificial adapter sequences added in the sequencing process. But the annotation of these tails is compl  ...[more]

Similar Datasets

| S-EPMC4713042 | biostudies-literature
| S-EPMC5272801 | biostudies-literature
| S-EPMC7159247 | biostudies-literature
| S-EPMC2532981 | biostudies-literature
| PRJEB11439 | ENA
| S-EPMC7173583 | biostudies-literature
| S-EPMC3493522 | biostudies-literature