Project description:While the importance of random sequencing errors decreases at higher DNA or RNA sequencing depths, systematic sequencing errors (SSEs) dominate at high sequencing depths and can be difficult to distinguish from biological variants. These SSEs can cause base quality scores to underestimate the probability of error at certain genomic positions, resulting in false positive variant calls, particularly in mixtures such as samples with RNA editing, tumors, circulating tumor cells, bacteria, mitochondrial heteroplasmy, or pooled DNA. Most algorithms proposed for correction of SSEs require a training data set, which is typically either from a part of the data set being “recalibrated” (Genome Analysis ToolKit, or GATK) or from a separate data set with special characteristics (SysCall). Here, we combine the advantages of these approaches by adding synthetic RNA spike-in standards to human RNA, and use GATK to recalibrate base quality scores with reads mapped to the spike-in standards. Compared to conventional GATK recalibration that uses reads mapped to the genome, spike-ins improve the accuracy of Illumina base quality scores by a mean of 5 units, and by as much as 13 units at CpG sites. In addition, since reads mapping to the genome are not used for recalibration, our method allows run-specific recalibration even for the many species without a comprehensive and accurate SNP database. We also use GATK with the spike-in standards to demonstrate that the Illumina RNA sequencing runs overestimate quality scores for AC, CC, GC, GG, and TC dinucleotides, while SOLiD has less dinucleotide SSEs but more SSEs for certain cycles. We conclude that using these DNA and RNA spike-in standards with GATK improves base quality score recalibration.
Project description:The purpose of this work was to describe a computational and analytical methodology for profiling small RNA by high-throughput sequencing. The datasets here were used to develop synthetic oligoribonucleotides as spike-in standards.
Project description:The purpose of this work was to describe a computational and analytical methodology for profiling small RNA by high-throughput sequencing. The datasets here were used to develop synthetic oligoribonucleotides as spike-in standards. We assessed the use of synthetic oligoribonucleotide standards as spike-in controls. These standards can be used to set an objective standard against which to compare samples. Standards were added to the total RNA (100 ug) in the following amounts: Std2 (TATATGCAAGTCCGGCCATAC) 0.01 pmol, Std3 (TAGCTAACGCATATCCGCATC) 0.1 pmol, Std6 (TGAAGCTGACATCGGTCATCC) 1.0 pmol.
Project description:While the importance of random sequencing errors decreases at higher DNA or RNA sequencing depths, systematic sequencing errors (SSEs) dominate at high sequencing depths and can be difficult to distinguish from biological variants. These SSEs can cause base quality scores to underestimate the probability of error at certain genomic positions, resulting in false positive variant calls, particularly in mixtures such as samples with RNA editing, tumors, circulating tumor cells, bacteria, mitochondrial heteroplasmy, or pooled DNA. Most algorithms proposed for correction of SSEs require a training data set, which is typically either from a part of the data set being M-bM-^@M-^\recalibratedM-bM-^@M-^] (Genome Analysis ToolKit, or GATK) or from a separate data set with special characteristics (SysCall). Here, we combine the advantages of these approaches by adding synthetic RNA spike-in standards to human RNA, and use GATK to recalibrate base quality scores with reads mapped to the spike-in standards. Compared to conventional GATK recalibration that uses reads mapped to the genome, spike-ins improve the accuracy of Illumina base quality scores by a mean of 5 units, and by as much as 13 units M-BM- at CpG sites. In addition, since reads mapping to the genome are not used for recalibration, our method allows run-specific recalibration even for the many species without a comprehensive and accurate SNP database. We also use GATK with the spike-in standards to demonstrate that the Illumina RNA sequencing runs overestimate quality scores for AC, CC, GC, GG, and TC dinucleotides, while SOLiD has less dinucleotide SSEs but more SSEs for certain cycles. We conclude that using these DNA and RNA spike-in standards with GATK improves base quality score recalibration. Four human RNA samples with equimolar ERCC spike-in standards were sequenced on Illumina. Two human brain/liver/muscle RNA mixtures with dynamic range of ERCC spike-in standards were sequenced on SOLiD.
Project description:Microarrays have become established tools for describing microbial systems, however the assessment of expression profiles for environmental microbial communities still presents unique challenges. Notably, the concentration of particular transcripts are likely very dilute relative to the pool of total RNA, and PCR-based amplification strategies are vulnerable to amplification biases and the appropriate primer selection. Thus, we apply a signal amplification approach, rather than template amplification, to analyze the expression of selected lignin-degrading enzymes in soil. Controls in the form of known amplicons and cDNA from Phanerochaete chrysosporium were included and mixed with the soil cDNA both before and after the signal amplification in order to assess the dynamic range of the microarray. We demonstrate that restored prairie soil expresses a diverse range of lignin-degrading enzymes following incubation with lignin substrate, while farmed agricultural soil does not. The mixed additions of control cDNA with soil cDNA indicate that the mixed biomass in the soil does interfere with low abundance transcript changes, nevertheless our microarray approach consistently reports the most robust signals. Keywords: comparative analysis, microbial ecology, soil microbial communities
Project description:In this work, we used a functional gene microarray approach (GeoChip) to assess the soil microbial community functional potential related to the different wine quality. In order to minimize the soil variability, this work was conducted at a “within-vineyard” scale, comparing two similar soils (BRO11 and BRO12) previously identified with respect to pedological and hydrological properties within a single vineyard in Central Tuscany and that yielded highly contrasting wine quality upon cultivation of the same Sangiovese cultivar