Project description:Genotyping of RpoD mutants via amplicon sequencing from the following manuscript: \\"Systematic dissection of σ70 sequence diversity and function in bacteria\\" by Park and Wang (2020). Includes raw sequencing reads from samples from MAGE-seq single codon saturation mutagenesis and high-throughput fitness competition experiment as well as the RpoD ortholog mutants generated through recombineering and CRISPR selection.
Project description:Purpose: The goal of this study is to compare endothelial small RNA transcriptome to identify the target of OASL under basal or stimulated conditions by utilizing miRNA-seq. Methods: Endothelial miRNA profilies of siCTL or siOASL transfected HUVECs were generated by illumina sequencing method, in duplicate. After sequencing, the raw sequence reads are filtered based on quality. The adapter sequences are also trimmed off the raw sequence reads. rRNA removed reads are sequentially aligned to reference genome (GRCh38) and miRNA prediction is performed by miRDeep2. Results: We identified known miRNA in species (miRDeep2) in the HUVECs transfected with siCTL or siOASL. The expression profile of mature miRNA is used to analyze differentially expressed miRNA(DE miRNA). Conclusions: Our study represents the first analysis of endothelial miRNA profiles affected by OASL knockdown with biologic replicates.
Project description:A cDNA library was constructed by Novogene (CA, USA) using a Small RNA Sample Pre Kit, and Illumina sequencing was conducted according to company workflow, using 20 million reads. Raw data were filtered for quality as determined by reads with a quality score > 5, reads containing N < 10%, no 5' primer contaminants, and reads with a 3' primer and insert tag. The 3' primer sequence was trimmed and reads with a poly A/T/G/C were removed
Project description:Whole exome sequencing of 5 HCLc tumor-germline pairs. Genomic DNA from HCLc tumor cells and T-cells for germline was used. Whole exome enrichment was performed with either Agilent SureSelect (50Mb, samples S3G/T, S5G/T, S9G/T) or Roche Nimblegen (44.1Mb, samples S4G/T and S6G/T). The resulting exome libraries were sequenced on the Illumina HiSeq platform with paired-end 100bp reads to an average depth of 120-134x. Bam files were generated using NovoalignMPI (v3.0) to align the raw fastq files to the reference genome sequence (hg19) and picard tools (v1.34) to flag duplicate reads (optical or pcr), unmapped reads, reads mapping to more than one location, and reads failing vendor QC.
Project description:The Caucasus, inhabited by modern humans since the Early Upper Paleolithic and known for its linguistic diversity, is considered to be important for understanding human dispersals and genetic diversity in Eurasia. We report a synthesis of autosomal, Y chromosome, and mitochondrial DNA (mtDNA) variation in populations from all major subregions and linguistic phyla of the area. Autosomal genome variation in the Caucasus reveals significant genetic uniformity among its ethnically and linguistically diverse populations and is consistent with predominantly Near/Middle Eastern origin of the Caucasians, with minor external impacts. In contrast to autosomal and mtDNA variation, signals of regional Y chromosome founder effects distinguish the eastern from western North Caucasians. Genetic discontinuity between the North Caucasus and the East European Plain contrasts with continuity through Anatolia and the Balkans, suggesting major routes of ancient gene flows and admixture.
Project description:Here, A549 cells expressing the ACE2 receptor were infected with SARS-CoV2, and pCHi-C was performed at 0 (mock), 8 and 24 hours post-infection. This repository provides the raw pCHi-C sequence reads and downstream processed CHiCAGO data (Rds files).
Project description:We report the sequences bound to CENP-A in the dog genome (Canis familiaris) for high-throughput characterization of centromeric sequences. We compare these ChIPSeq reads (72 bp, single read) against a reference centromeric satellite DNA domain database for the dog genome, resulting in the annotation of sequence variation and estimated abundance of seven satellite families together with adjacent, non-satellite sequences. To study global patterns of sequence diversity and characterizing the subset of sequences correlated with centromere function, these sequences were evaluated relative to a comprehensive centromere sequence domain k-mer library. From this analysis, we identify functional sequence features from two satellite families (CarSat1 and CarSat2) that are defined by distinct arrays subtypes. Sequences bound to CENP-A in MDCK (dog) cell line
Project description:HDMYZ cells were treated with 2ug/ml ActD for 0, 4 and 12 hours. Small RNAs of 15-40 bases were gel-purified from 10 ug total RNA, and subjected to multiplex Illumina small RNA library preparation. Small RNA libraries were sequenced on a HiSeq2000 (Illumina) with 3 samples per lane. To quantify miRNA and isoform abundance, sequence reads were processed by the miRDeep2 package, with the following modifications. First, to remove adaptor sequence, we removed both the main adaptor sequence present in the sequencing reads, as well as the second most abundant adaptor variant. In addition, we did not restrict the size of small RNAs during adaptor removal. Second, we used miRBase v18 for mapping the reads. Third, for quantifying miRNA and isoform frequency, we limited reads to more or equal to 15 bases in length with zero mis-match during mapping. The number of reads that were mapped to known miRNAs was used to normalize read frequencies for each miRNA or each miRNA isoform. For quantification purposes, we only considered miRNAs or isoforms that had frequency >= 1x10e-6 in samples without ActD treatment, which correspond to ~21-30 reads in raw count. These miRNAs or isoforms were referred to as reliably quantifiable.To analyze mapping to the genome, we removed reads that mapped to miRNA precursors. The rest of the reads were then mapped to the genome with Bowtie.