GNPS - Chemical Standards - Bile Acids and Conjugates, Human
Ontology highlight
ABSTRACT: data from the analysis of chemical standards analyzed in positive and negative ionization mode. data collected using data-dependent acquisition.
Project description:MS/MS fragmentation data on bile acid standards were acquired on the QE - with a gradient developed to separate between isomeric pairs on a Polar C18 column and a fragmentation energy of NCE 45.
Project description:LC-MS/MS data were collected from chemical standards including 1-Dehydroandrostenedion (5uM), Chenodeoxyglycocholic acid (5uM), Dehydroisoandrosterone sulfate (100uM), Glycocholic acid (5uM) and Taurocholic acid (5uM).
Project description:While the importance of random sequencing errors decreases at higher DNA or RNA sequencing depths, systematic sequencing errors (SSEs) dominate at high sequencing depths and can be difficult to distinguish from biological variants. These SSEs can cause base quality scores to underestimate the probability of error at certain genomic positions, resulting in false positive variant calls, particularly in mixtures such as samples with RNA editing, tumors, circulating tumor cells, bacteria, mitochondrial heteroplasmy, or pooled DNA. Most algorithms proposed for correction of SSEs require a training data set, which is typically either from a part of the data set being “recalibrated” (Genome Analysis ToolKit, or GATK) or from a separate data set with special characteristics (SysCall). Here, we combine the advantages of these approaches by adding synthetic RNA spike-in standards to human RNA, and use GATK to recalibrate base quality scores with reads mapped to the spike-in standards. Compared to conventional GATK recalibration that uses reads mapped to the genome, spike-ins improve the accuracy of Illumina base quality scores by a mean of 5 units, and by as much as 13 units at CpG sites. In addition, since reads mapping to the genome are not used for recalibration, our method allows run-specific recalibration even for the many species without a comprehensive and accurate SNP database. We also use GATK with the spike-in standards to demonstrate that the Illumina RNA sequencing runs overestimate quality scores for AC, CC, GC, GG, and TC dinucleotides, while SOLiD has less dinucleotide SSEs but more SSEs for certain cycles. We conclude that using these DNA and RNA spike-in standards with GATK improves base quality score recalibration.
Project description:Untargeted UPLC-MS/MS data from a screen of human gut microbiome commensal isolate bacteria from healthy human donor feces. Isolates include Mediterraneibacter gnavus MSK15.77 (NCBI accession NZ_JAAIRR010000000), Bacteroides ovatus MSK22.29 (NCBI accession NZ_JAHOCX010000000), Bifidobacterium longum DFI.2.45 (NCBI accession NZ_JAJCNS010000000), and Lachnoclostridium scindens SL.1.22 (NCBI accession GCA_020555615.1). The dataset includes validated bile acid standards. All using positive ionization.
Project description:While the importance of random sequencing errors decreases at higher DNA or RNA sequencing depths, systematic sequencing errors (SSEs) dominate at high sequencing depths and can be difficult to distinguish from biological variants. These SSEs can cause base quality scores to underestimate the probability of error at certain genomic positions, resulting in false positive variant calls, particularly in mixtures such as samples with RNA editing, tumors, circulating tumor cells, bacteria, mitochondrial heteroplasmy, or pooled DNA. Most algorithms proposed for correction of SSEs require a training data set, which is typically either from a part of the data set being M-bM-^@M-^\recalibratedM-bM-^@M-^] (Genome Analysis ToolKit, or GATK) or from a separate data set with special characteristics (SysCall). Here, we combine the advantages of these approaches by adding synthetic RNA spike-in standards to human RNA, and use GATK to recalibrate base quality scores with reads mapped to the spike-in standards. Compared to conventional GATK recalibration that uses reads mapped to the genome, spike-ins improve the accuracy of Illumina base quality scores by a mean of 5 units, and by as much as 13 units M-BM- at CpG sites. In addition, since reads mapping to the genome are not used for recalibration, our method allows run-specific recalibration even for the many species without a comprehensive and accurate SNP database. We also use GATK with the spike-in standards to demonstrate that the Illumina RNA sequencing runs overestimate quality scores for AC, CC, GC, GG, and TC dinucleotides, while SOLiD has less dinucleotide SSEs but more SSEs for certain cycles. We conclude that using these DNA and RNA spike-in standards with GATK improves base quality score recalibration. Four human RNA samples with equimolar ERCC spike-in standards were sequenced on Illumina. Two human brain/liver/muscle RNA mixtures with dynamic range of ERCC spike-in standards were sequenced on SOLiD.
Project description:<p>To better understand the effect of different LC-MS setups on metabolite detection, five laboratories analyzed 1298 standard compounds obtained from the MetaSci metabolite standard library, classified as 'Human Endosome' (HE), 'Food Exposome' (FE), 'Chemical Exposome' (CE), 'Microbiota Exposome' (ME), 'Plant Exposome' (PE), and 'Eukaryotic Exposome' (EE), using a common reference reverse-phase method in both positive and negative ionization modes and different in-house methods. NAPS were used for alignment and retention time indexing.</p><p>In our laboratory, we ran the reference reverse-phase method, an in-house reverse-phase method, and an in-house HILIC method, all three with positive and negative ionization, using an LC-MS with a Q-ToF analyzer. In this study, we present data from the reference reverse-phase method and the in-house reverse-phase method.</p>