Project description:Deregulated gene expression is a hallmark of cancer, however most studies to date have analyzed short-read RNA-sequencing data with inherent limitations. Here, we combine PacBio long-read isoform sequencing (Iso-Seq) and Illumina paired-end short read RNA sequencing to comprehensively survey the transcriptome of gastric cancer (GC), a leading cause of global cancer mortality. We performed full-length transcriptome analysis across 10 GC cell lines covering four major GC molecular subtypes (chromosomal unstable, Epstein-Barr positive, genome stable and microsatellite unstable). We identify 60,239 non-redundant full-length transcripts, of which >66% are novel compared to current transcriptome databases. Novel isoforms are more likely to be cell-line and subtype specific, expressed at lower levels with larger number of exons, with longer isoform/coding sequence lengths. Most novel isoforms utilize an alternate first exon, and compared to other alternative splicing categories are expressed at higher levels and exhibit higher variability. Collectively, we observe alternate promoter usage in 25% of detected genes, with the majority (84.2%) of known/novel promoter pairs exhibiting potential changes in their coding sequences. Mapping these alternate promoters to TCGA GC samples, we identify several cancer-associated isoforms, including novel variants of oncogenes. Tumor-specific transcript isoforms tend to alter protein coding sequences to a larger extent than other isoforms. Analysis of outcome data suggests that novel isoforms may impart additional prognostic information. Our results provide a rich resource of full-length transcriptome data for deeper studies of GC and other gastrointestinal malignancies.
Project description:Clinically translatable large animal models have become indispensable for cardiovascular research, clinically relevant proof of concept studies and for novel diagnostic and therapeutic interventions. In particular, the pig as emerged as an essential cardiovascular disease model, because its heart, circulatory system, and blood supply are anatomically and functionally similar to that of humans. Unfortunately, molecular and omics-based studies in the pig are hampered by the incompleteness of the genome and the lack of diversity of the corresponding transcriptome annotation. Here, we employed Nanopore long-read sequencing and in-depth proteomics on top of Illumina RNA-seq to enhance the pig cardiac transcriptome annotation. We assembled 15,926 transcripts, stratified into coding and non-coding, and validated our results by complementary mass spectrometry. A manual review of several gene loci, which are associated with cardiac function, corroborated the utility of our enhanced annotation. All our data are available for download and is also provided as tracks for integration in genome browsers. We deem this resource as highly valuable for molecular research in an increasingly relevant large animal model.
Project description:The primary objective of this prospective observational study is to characterize the gut and oral microbiome as well as the whole blood transcriptome in gastrointestinal cancer patients and correlate these findings with cancer type, treatment efficacy and toxicity. Participants will be recruited from existing clinical sites only, no additional clinical sites are needed.
Project description:Tanaidaceans are small benthic crustaceans that mainly inhabit diverse marine environments, and they comprise one of the most diverse and abundant macrofaunal groups in the deep sea. Tanaidacea is one of the most thread-dependent taxa in the Crustacea, constructing tube spun with their silk for shelter. In this work, we sequenced and assembled the comprehensive transcriptome of 23 tanaidaceans encompassing 14 families and 4 superfamilies of Tanaidacea, and performed silk proteomics of Zeuxo ezoensis to search for its silk genes. As a result, we identified two families of silk proteins, that are conserved across the four superfamilies. Long and repetitive nature of these silk genes resemble that of other silk-producing organisms, and the two families of proteins were similar in composition to silkworm and caddisform fibroins, respectively. Moreover, the amino acid composition of the repetitive motifs of tanaidid silk tended to be more hydrophilic, and therefore could be a useful resource to study their unique adaptation of silk use in marine environment. The availability of comprehensive transcriptome data in these taxa coupled with the proteomics evidence for their silk genes would facilitate the evolutionary and ecological studies.
Project description:Microbiome sequencing model is a Named Entity Recognition (NER) model that identifies and annotates microbiome nucleic acid sequencing method or platform in texts. This is the final model version used to annotate metagenomics publications in Europe PMC and enrich metagenomics studies in MGnify with sequencing metadata from literature. For more information, please refer to the following blogs: http://blog.europepmc.org/2020/11/europe-pmc-publications-metagenomics-annotations.html https://www.ebi.ac.uk/about/news/service-news/enriched-metadata-fields-mgnify-based-text-mining-associated-publications
Project description:We produced an extensive transcript catalog for LCLs of 5 primate species by leveraging isoform sequencing and short-read RNA-seq. The curated transcriptomes were used to assist mass spectrometry protein identifications.
Project description:In Europe, ticks are the most important vectors of diseases threatening humans, livestock, wildlife and companion animals. Nevertheless, genomic sequence information and functional annotation of proteins of the most important European tick, Ixodes ricinus, is limited. Here we present the first analysis of the I. ricinus genome and of the transcriptome of the unfed I. ricinus midgut. We combined and integrated data from genome, transcriptome and proteome. The de novo assembly of 1 billion paired-end sequences identified 6,415 putative genes providing an unprecedented insight into the I. ricinus genome. Mapping of our midgut mRNA reads to the assembled contigs let us estimate to cover around two third of the unique genomic sequences. In addition, more than 10,000 transcripts from naïve midgut were annotated functionally and/or locally. By combining the alignment-based with a motif-search based annotation approach, we could double the number of annotations throughout all groups without shifting the dataset. Moreover, 1,175 proteins expressed in the naïve midgut were identified by mass spectrometry confirming the high completeness of our transcriptome database, and 608 were significantly annotated for function and/or localization. This multiple-omics study vastly extends the publicly available DNA, RNA and protein databases for I. ricinus and ticks in general.
Project description:This SuperSeries is composed of the following subset Series: GSE34461: Comparing two transcriptome technologies - sequencing match microarrays [Array] GSE34477: Comparing two transcriptome technologies - sequencing match microarrays [RNA-Seq] Refer to individual Series
Project description:Endothelial cell (EC) metabolism is an emerging target for anti-angiogenic therapy in tumor and choroidal neovascularization (CNV), but little is known about individual EC metabolic transcriptomes. Here, by scRNA-sequencing 28,337 murine choroidal ECs (CECs) and sprouting CNV-ECs, we constructed a taxonomy to characterize their heterogeneity. Comparison with murine lung tumor ECs (TECs) revealed congruent marker gene expression by distinct EC phenotypes across tissues and diseases, suggesting similar angiogenic mechanisms. Trajectory inference of CNV-ECs revealed that differentiation of venous to angiogenic ECs was accompanied by metabolic transcriptome plasticity. EC phenotypes displayed metabolic transcriptome heterogeneity. Hypothesizing that conserved genes are more important, we used an integrated analysis, based on congruent transcriptome analysis, CEC-tailored genome scale metabolic modeling, and gene expression meta-analysis in multiple cross-species datasets, followed by functional validation, to identify the top-ranking metabolic targets SQLE and ALDH18A1, involved in EC proliferation and collagen production, respectively, as novel angiogenic targets. The effect of SQLE and ALDH18A1 silencing in ECs was investigated by transcriptomics and proteomics analysis.