Project description:Two PacBio Hifi sequencing runs from the kidney of a single male NMR sample used to make the mHetGla4.1.primary genome assembly. Specifically, we assembled a second NMR genome from an unrelated male of a separate captive colony in Toronto, Canada, using PacBio HiFi (155.8 Gb, read N50 = 11.45 Kb) and ONT-LSK (299.6 Gb, read N50 = 10.1 Kb) reads (contig N50 = 75.7 Mb, Compleasm S = 98%). This accession stores the ONT-LSK data for this independent assembly.
Project description:Two PacBio Hifi sequencing runs from the kidney of a single male NMR sample used to make the mHetGla4.1.primary genome assembly. Specifically, we assembled a second NMR genome from an unrelated male of a separate captive colony in Toronto, Canada, using PacBio HiFi (155.8 Gb, read N50 = 11.45 Kb) and ONT-LSK (299.6 Gb, read N50 = 10.1 Kb) reads (contig N50 = 75.7 Mb, Compleasm S = 98%). This accession stores the Pacbio Hifi data for this independent assembly.
Project description:One PacBio Hifi sequencing run from the kidney of a single male CDMR (Bathyergus suillus) sample used to make the mBatSui1.1.primary genome assembly as an evolutionary comparator to our telomere-to-telomere naked mole-rat genome assembly. Specifically, we assembled a CDMR from a wild-derived sample in South African cape and sequenced in Toronto, Canada, using PacBio HiFi (89 Gb, read N50 = 18 Kb) and ONT-ULK (55 Gb, read N50 = 43 Kb) reads (contig N50 = 33 Mb, Compleasm S = 99%, QV = 71.0). This accession stores the Pacbio Hifi data for this assembly.
Project description:One ONT-ULK sequencing run from the kidney of a single male CDMR (Bathyergus suillus) sample used to make the mBatSui1.1.primary genome assembly as an evolutionary comparator to our telomere-to-telomere naked mole-rat genome assembly. Specifically, we assembled a CDMR from a wild-derived sample in South African cape and sequenced in Toronto, Canada, using PacBio HiFi (89 Gb, read N50 = 18 Kb) and ONT-ULK (55 Gb, read N50 = 43 Kb) reads (contig N50 = 33 Mb, Compleasm S = 99%, QV = 71.0). This accession stores the ONT-ULK data for this assembly.
Project description:long-read CAGE was design to identify full length capped transcript across 10 specific loci in cortical neurones. Long-read CAGE was based on the Cap-Trapper method with the full length cDNA sequencing using ONT MinION sequencer. After RNA extraction, 10 µg total RNAs from Human iPS (WTC-11) cells, differentiated neural stem cells and differentiated cortical neuron cells were polyadenylated with E-coli poly(A) Polymerase (PAP) (NEB M0276) at 37°C for 15 min and purified with AMPure RNA Clean XP beads. The PAP treated 5 µg RNA was reverse transcribed with oligodT_16VN_UMI25_primer (GAGATGTCTCGTGGGCTCGGNNNNNNNNNNNNNNNNNNNNNNNNNCTACGTTTTTTTTTTTTTTTTVN) and Prime Script II Reverse Transcriptase (Takara Bio) at 42°C for 60 min and purified with RNAClean XP beads. Cap-trapping from the RNA/cDNA hybrids was performed with published protocol (Takahashi et al., Nature protocols, 2012 (https://doi.org/10.1038/nprot.2012.005)), and RNA was digested with RNase H (Takara Bio) at 37°C for 30 min and purified with AMPureXP beads. 5’ linker (N6 up GTGGTATCAACGCAGAGTACNNNNNN-Phos, GN5 up GTGGTATCAACGCAGAGTACGNNNNN-Phos, down Phos-GTACTCTGCGTTGATACCAC-Phos) was ligated to the cDNA with Mighty Mix (Takara Bio) for overnight and the ligated cDNA was purified with AMPure XP beads. Shrimp Alkaline Phosphatase (Takara Bio) was used to remove phosphates at the ligated linker and purified with AMPureXP beads. The 5’ linker ligated cDNA was then second strand synthesized with KAPA HiFi mix (Roche) and 2nd synthesis primer_UMI15 at 95°C for 5 min, 55°C for 5 min and 72°C for 30 min. Exonuclease I (Takara Bio) was added for the primer digestion at 37°C for 30 min, and the cDNA/DNA hybrid was purified with AMPureXP and amplified with PrimerSTAR GXL DNA polymerase (Takara Bio) and PCR primer (fwd_CTACACTCGTCGGCAGCGTC, rev _GAGATGTCTCGTGGGCTCGG) for 7 cycles. The library was then treated with SQK-LSK110 (Oxford Nanopore Technologies) with manufacture’s protocol and sequenced with R9.4 flowcell (FLO-MIN106) in MinION sequencer. Basecalling was processed by Guppy v5.0.14 basecaller software provided by Oxford Nanopore Technologies to generate fastq files from FAST5 files. To prepare clean reads from fastq files, adapter sequence was trimmed by pychopper (https://github.com/nanoporetech/pychopper) with VNP_GAGATGTCTCGTGGGCTCGGNNNNNNNNNNNNNNNCTACG and SSP_ CTACACTCGTCGGCAGCGTCNNNNNNNNNNNNNNNNNNNNNNNNNGTGGTATCAACGCAGAGTAC and the fastq was mapped on our target genes.
Project description:BmN4 cells are cultured cells derived from Bombyx mori ovaries and widely used to study transposon silencing by PIWI-interacting RNAs (piRNAs). A high-accurate genome sequence of BmN4 cells is required to analyze the piRNA pathway using RNA-seq. The genome sequence of BmN4 cells was assembled using Pacific Biosciences (PacBio) HiFi and Oxford Nanopore technology Ultralong (ONT-UL) reads. Microscopic observation and image analysis showed that BmN4 cells were octoploid on average, and the number of chromosomes per cell was highly variable. We concluded the haplotype-resolved assembly of such a complex genome would be difficult; therefore, we assembled a consensus genome sequence. RNA-seq analysis of Siwi knockdown cells also revealed that Siwi-piRISC may target Countdown (Cd), an LTR retrotransposon. By comparing the consensus genome sequence with the reads, we identified differences between haplotypes, particulary structural variants, suggesting that some transposons, including Countdown, increased their copy number in BmN4 cells.
Project description:Alternative splicing is widely acknowledged to be a crucial regulator of gene expression and is a key contributor to both normal developmental processes and disease states. While cost-effective and accurate for quantification, short-read RNA-seq lacks the ability to resolve full-length transcript isoforms despite increasingly sophisticated computational methods. Long-read sequencing platforms such as Pacific Biosciences (PacBio) and Oxford Nanopore (ONT) bypass the transcript reconstruction challenges of short-reads. Here we describe TALON, the ENCODE4 pipeline for analyzing PacBio cDNA and ONT direct-RNA transcriptomes. We apply TALON to three human ENCODE Tier 1 cell lines and show that while both technologies perform well at full-transcript discovery and quantification, each technology has its distinct artifacts. We further apply TALON to mouse cortical and hippocampal transcriptomes and find that a substantial proportion of neuronal genes have more reads associated with novel isoforms than annotated ones. The TALON pipeline for technology-agnostic, long-read transcriptome discovery and quantification tracks both known and novel transcript models as well as expression levels across datasets for both simple studies and larger projects such as ENCODE that seek to decode transcriptional regulation in the human and mouse genomes to predict more accurate expression levels of genes and transcripts than possible with short-reads alone.