Accurate identification and analysis of human mRNA isoforms using deep long read sequencing.
Ontology highlight
ABSTRACT: Precise identification of RNA-coding regions and transcriptomes of eukaryotes is a significant problem in biology. Currently, eukaryote transcriptomes are analyzed using deep short-read sequencing experiments of complementary DNAs. The resulting short-reads are then aligned against a genome and annotated junctions to infer biological meaning. Here we use long-read complementary DNA datasets for the analysis of a eukaryotic transcriptome and generate two large datasets in the human K562 and HeLa S3 cell lines. Both data sets comprised at least 4 million reads and had median read lengths greater than 500 bp. We show that annotation-independent alignments of these reads provide partial gene structures that are very much in-line with annotated gene structures, 15% of which have not been obtain
SUBMITTER: Tilgner H
PROVIDER: S-EPMC3583448 | biostudies-literature | 2013 Mar
REPOSITORIES: biostudies-literature
ACCESS DATA