Unknown

Dataset Information

0

Contrasting and combining transcriptome complexity captured by short and long RNA sequencing reads.


ABSTRACT: Mapping transcriptomic variations using either short- or long-read RNA sequencing is a staple of genomic research. Long reads are able to capture entire isoforms and overcome repetitive regions, whereas short reads still provide improved coverage and error rates. Yet, open questions remain, such as how to quantitatively compare the technologies, can we combine them, and what is the benefit of such a combined view? We tackle these questions by first creating a pipeline to assess matched long- and short-read data using a variety of transcriptome statistics. We find that across data sets, algorithms, and technologies, matched short-read data detects ∼30% more splice junctions, such that ∼10%-30% of the splice junctions included at ≥20% by short reads are missed by long reads. In contrast, long reads detect many more intron-retention events and can detect full isoforms, pointing to the benefit of combining the technologies. We introduce MAJIQ-L, an extension of the MAJIQ software, to enable a unified view of transcriptome variations from both technologies and demonstrate its benefits. Our software can be used to assess any future long-read technology or algorithm and can be combined with short-read data for improved transcriptome analysis.

SUBMITTER: Han SW 

PROVIDER: S-EPMC11529863 | biostudies-literature | 2024 Oct

REPOSITORIES: biostudies-literature

altmetric image

Publications

Contrasting and combining transcriptome complexity captured by short and long RNA sequencing reads.

Han Seong Woo SW   Jewell San S   Thomas-Tikhonenko Andrei A   Barash Yoseph Y  

Genome research 20241029 10


Mapping transcriptomic variations using either short- or long-read RNA sequencing is a staple of genomic research. Long reads are able to capture entire isoforms and overcome repetitive regions, whereas short reads still provide improved coverage and error rates. Yet, open questions remain, such as how to quantitatively compare the technologies, can we combine them, and what is the benefit of such a combined view? We tackle these questions by first creating a pipeline to assess matched long- and  ...[more]

Similar Datasets

| S-EPMC10690182 | biostudies-literature
| S-EPMC6822431 | biostudies-literature
| S-EPMC8665758 | biostudies-literature
| S-EPMC7584020 | biostudies-literature
2024-07-10 | GSE271528 | GEO