Project description:Durum wheat (Triticum turgidum L. ssp. durum) is a major cereal and staple in the semi-arid regions of the Mediterranean Basin. It originates from BBAA wild tetraploid domesticated in Neolithic era, later evolving to domesticated emmer and then to up to 11 T. turgidum subspecies, including durum wheat landraces and modern cultivars. Tetraploid wheat is the donor of the A and B genomes of hexaploid bread wheat (DDAABB), representing therefore a valuable source of genetic variability and beneficial alleles for both durum and bread wheat breeding. After assembling the Platinum-quality reference genome for Svevo durum wheat cultivar coupling PACBIO HiFi long read 35X sequencing with BIONANO Optical Mapping and Hi-C conformation capture, a complete and accurate gene annotation was then obtained by coupling Illumina RNASeq and Nanopore Isoseq sequencing from multiple tissues. The expression of 68,154 high confidence genes together with more than 100,000 low confidence, TE-related or long non-coding genes was investigated on 30 diverse tissues from grain, root, leaf, and spike samples across multiple developmental time points to create a transcriptional atlas of durum wheat development.
Project description:Accurate annotations of genes and their transcripts is a foundation of genomics, but no annotation technique presently combines throughput and accuracy. As a result, the GENCODE reference collection of long noncoding RNAs remains far from complete: many are fragmentary, while thousands more remain uncatalogued. To accelerate lncRNA annotation, we have developed RNA Capture Long Seq (CLS), combining targeted RNA capture with third generation long-read sequencing. We present an experimental re-annotation of the entire GENCODE intergenic lncRNA populations in matched human and mouse tissues. CLS approximately doubles the complexity of targeted loci, both in terms of validated splice junctions and transcript models. Through its identification of full-length transcript models, CLS allows the first definitive measurement of promoter features, gene structure and protein-coding potential of lncRNAs. Thus CLS removes a longstanding bottleneck of transcriptome annotation, generating manual-quality full-length transcript models at high-throughput scales.
Project description:This dataset provides transcriptome sequencing data from a cultured female Quasipaa spinosa individual collected in Taining, Fujian Province, China (Q. spinosa-FJTN). The transcriptomic data were generated to support genome annotation of a newly assembled chromosome-scale reference genome. Ten adult tissues, including brain, heart, liver, kidney, lung, intestine, skin, spleen, ovary, and testis, were sampled, and total RNA was extracted separately from each tissue. Equal amounts of high-quality RNA from all tissues were pooled to construct a comprehensive RNA sequencing library. The resulting transcriptome dataset was generated using the DNBSEQ sequencing platform and provides expressed sequence evidence for protein-coding gene prediction and functional annotation. These data complement long-read genome sequencing resources, including PacBio HiFi, Oxford Nanopore ultra-long reads, and Hi-C sequencing, and contribute to the development of genomic resources for Q. spinosa population genetics, comparative genomics, conservation studies, and selective breeding.
Project description:Understanding gene expression diversity across human populations is essential for accurate genome annotation and disease interpretation. However, existing annotations are primarily based on European-derived transcriptomic data, potentially limiting their applicability to other populations. This study aims to assess population-specific transcript diversity and its impact on gene annotation. To achieve this, we performed long-read RNA sequencing on lymphoblastoid cell lines from 43 individuals across eight globally diverse populations. Our workflow included RNA extraction, cDNA synthesis, and sequencing using Oxford Nanopore long-read technology, followed by transcript assembly and comparison with existing gene annotations. We also integrated novel transcripts into reference annotations to evaluate their effect on allele-specific transcript usage detection. This work provides a critical step toward improving transcriptome annotation across diverse populations, ensuring a more comprehensive representation of human genetic variation. We provide here unprocessed, unaligned BAMs (just basecalled; uploaded file names end with *.bam) along with (ONT duplex read resolution, chimeric read splitting, UMI-based deduplication, adapter trimming, and basecalling quality >= 10; uploaded file names end with *preprocessed_Q10.fastq.gz) for each sample.