Project description:Chromosomal copy number variations (CNV) have been associated with various neurological and developmental disorders and chromosomal microarray (CMA) is a method of choice to diagnose Copy Number Gain/Loss syndromes. Recently, next-generation sequencing (NGS)-based low-coverage whole genome sequencing (LC-WGS) has been applied to detect Copy Number Gain/Loss syndromes. This dataset is intended to be used as a “Golden standard data set” for development of LC-WGS analysis method. It consists of patients (n=63) who have a mental delay and/or physical disability phenotype and normal (n=20) phenotype.
2018-09-29 | GSE120624 | GEO
Project description:Low coverage sequencing data of Duroc boars
| PRJEB44569 | ENA
Project description:Low coverage sequencing data of Duroc boars
Project description:Embryonic genome activation (EGA), a pivotal transcriptional event during preimplantation development, is accompanied by post-transcriptional regulation of maternal mRNAs. Disentangling the transcriptional output of the newly activated embryonic genome from concomitant post-transcriptional processing is important for decoding EGA dynamics.Here, using optimized low-input SLAM-seq (thiol(SH)-linked alkylation for the metabolic sequencing) in mouse embryos, we delineates the temporal hierarchy of EGA nascent transcription during mouse preimplantation embryogenesis and uncovers a mechanistic link between EGA and the first lineage specification, providing new insights into the regulatory architecture of early mammalian development.
Project description:In order to validate of CNV detection from low-coverage whole-genome sequencing in the blood samples from recurrent miscarriage couples, we employed a customized array Comparative Genomics Hybridization (aCGH, Agilent) approach as chromosomal microarray analysis (CMA) in present study for a cohort of 78 DNA samples from blood. CMA results were compared with low-coverage whole-genome sequencing detection results. 100% consistency was obtained in pathogenic or likely pathogenic CNVs detection.
Project description:Next generation sequencing platforms have become essential tools for understanding DNA in a wide range of contexts. Their success heavily relies on the accuracy, sensitivity and specificity of methods used to discern differences between the reference genome and genomes under investigation. Here we compare the relative performances of five popular single nucleotide variant callers with and without their associated recommended hard filtering criteria. We compare: FreeBayes; the Genome Analysis Toolkit’s Haplotype Caller and Unified Genotyper; SAMtools; and VarScan. We tailor this comparison to suit smaller projects with modest sample numbers (n = 10) and coverage (~10X) to fill a current gap in the literature. Other comparison studies are generally applicable only to larger projects in model species, where there is access to large amounts of sequencing data and curated callsets for base and variant quality score recalibration. We estimated the accuracy, sensitivity and specificity of each pipeline according to the genotype concordance rate and number with the “truth” dataset for 10 canine samples. The truth dataset was defined as genotypes obtained from the CanineHD BeadChip array. Whole genome sequencing data was performed on the Illumina HiSeq2000 or HiSeq2500 platform as 100-101 base pair, paired end reads to an average sample coverage of 10.3X. Apart from GATK Haplotype Caller, applying recommended hard filters did not improve the performance of genotyping concordance at the tested levels of minimum coverage. The default VarScan pipeline with no additional filters applied (VarScan uses SAMtools mpileup, without base alignment quality computation) generally outperformed other callers in terms of accuracy, sensitivity and specificity. The results of this study demonstrate that hard filtering of variant calls from low-powered genome studies can impair accuracy, sensitivity and specificity of callsets and provides some benchmark performance metrics on a range of low coverage levels.
Project description:Mass spectrometry (MS)-based proteomics aims to characterize comprehensive proteomes in a fast and reproducible manner. Here, we present an ultra-fast scanning data-independent acquisition (DIA) strategy consisting on 2-Th precursor isolation windows, dissolving the differences between data-dependent and independent methods. This is achieved by pairing a Quadrupole Orbitrap mass spectrometer with the asymmetric track lossless (Astral) analyzer that provides >200 Hz MS/MS scanning speed, high resolving power and sensitivity, as well as low ppm-mass accuracy. Narrowwindow DIA enables profiling of up to 100 full yeast proteomes per day, or ~10,000 human proteins in half-an-hour. Moreover, multi-shot acquisition of fractionated samples allows comprehensive coverage of human proteomes in ~3h, showing comparable depth to next-generation RNA sequencing and with 10x higher throughput compared to current state-of-the-art MS. High quantitative precision and accuracy is demonstrated with high peptide coverage in a 3-species proteome mixture, quantifying 14,000+ proteins in a single run in half-an-hour.
Project description:This experiment contains a subset of data from the BLUEPRINT Epigenome project ( http://www.blueprint-epigenome.eu ), which aims at producing a reference haemopoetic epigenomes for the research community. 74 samples of primary cells or cultured primary cells of different haemopoeitc lineages from cord blood, venous blood, bone marrow and thymus are included in this experiment. This ArrayExpress record contains only meta-data. Raw data files have been archived at the European Genome-Phenome Archive (EGA, www.ebi.ac.uk/ega) by the consortium, with restricted access to protect sample donors' identity. There are 32 EGA data set accessions, which can be found under the Comment[EGA_DATA_SET] column in the 'Sample Data Relationship Format' (SDRF) file of this ArrayExpress record (http://www.ebi.ac.uk/arrayexpress/files/E-MTAB-3827/E-MTAB-3827.sdrf.txt). Details on how to apply for data access via the BLUEPRINT data access committee are on the EGA data set pages. Likewise, mapping of samples to these EGA accessions can be found in the SDRF file. Please note that the raw data files for 11 sequencing runs have yet been deposited at EGA, so they are marked with \\ot available\\ under the Comment[SUBMITTED_FILE_NAME] field in the SDRF file, and were included for the sake of completeness. Further iInformation on individual samples and sequencing libraries can also be found on the BLUEPRINT data coordination centre (DCC) website: http://dcc.blueprint-epigenome.eu\
Project description:This experiment contains a subset of data from the BLUEPRINT Epigenome project ( http://www.blueprint-epigenome.eu ), which aims at producing a reference haemopoetic epigenomes for the research community. 4 samples of primary cells from tonsil with cell surface markes CD20med/CD38high in young individuals (3 to 10 years old) are included in this experiment. This ArrayExpress record contains only meta-data. Raw data files have been archived at the European Genome-Phenome Archive (EGA, www.ebi.ac.uk/ega) by the consortium, with restricted access to protect sample donors' identity. The relevant accessions of EGA data sets is EGAD00001001523. Details on how to apply for data access via the BLUEPRINT data access committee are on the EGA data set pages. The mapping of samples to these EGA accessions can be found in the 'Sample Data Relationship Format' file of this ArrayExpress record. Information on individual samples and sequencing libraries can also be found on the BLUEPRINT data coordination centre (DCC) website: http://dcc.blueprint-epigenome.eu
Project description:This experiment contains a subset of data from the BLUEPRINT Epigenome project ( http://www.blueprint-epigenome.eu ), which aims at producing a reference haemopoetic epigenomes for the research community. 29 samples of primary cells or cultured primary cells of different haemopoeitc lineages from cord blood are included in this experiment. This ArrayExpress record contains only meta-data. Raw data files have been archived at the European Genome-Phenome Archive (EGA, www.ebi.ac.uk/ega) by the consortium, with restricted access to protect sample donors' identity. The relevant accessions of EGA data sets is EGAD00001001165. Details on how to apply for data access via the BLUEPRINT data access committee are on the EGA data set pages. The mapping of samples to these EGA accessions can be found in the 'Sample Data Relationship Format' file of this ArrayExpress record. Information on individual samples and sequencing libraries can also be found on the BLUEPRINT data coordination centre (DCC) website: http://dcc.blueprint-epigenome.eu