Genomics

Dataset Information

0

Accurate annotation of human protein-coding small open reading frames


ABSTRACT: Protein-coding small open reading frames (smORFs) are emerging as an important class of genes, however, the coding capacity of smORFs in the human genome is unclear. By integrating de novo transcriptome assembly and Ribo-Seq, we confidently annotate thousands of novel translated smORFs in three human cell lines. We find that smORF translation prediction is noisier than for annotated coding sequences, underscoring the importance of analyzing multiple experiments and footprinting conditions. These smORFs are located within non-coding and antisense transcripts, the UTRs of mRNAs, and unannotated transcripts. Analysis of RNA levels and translation efficiency during cellular stress identifies regulated smORFs and provides an approach for identifying smORFs for further investigation. Sequence conservation and signatures of positive selection indicate that encoded microproteins are likely functional. Additionally, proteomics data from enriched human leukocyte antigen complexes validates the translation of hundreds of smORFs and positions them as a source of novel antigens. Thus, smORFs represent a significant number of important, yet unexplored human genes.

ORGANISM(S): Homo sapiens

PROVIDER: GSE125218 | GEO | 2019/07/03

REPOSITORIES: GEO

Similar Datasets

2014-09-11 | E-GEOD-60384 | biostudies-arrayexpress
2020-03-14 | GSE131650 | GEO
2021-04-28 | GSE154491 | GEO
2017-10-17 | PXD005643 | Pride
2022-12-13 | GSE198107 | GEO
2022-12-13 | GSE197909 | GEO
2019-06-26 | MSV000084014 | MassIVE
2014-09-11 | GSE60384 | GEO
2017-12-20 | GSE92659 | GEO
2018-04-30 | GSE105082 | GEO