Project description:Chromosomal copy number variations (CNV) have been associated with various neurological and developmental disorders and chromosomal microarray (CMA) is a method of choice to diagnose Copy Number Gain/Loss syndromes. Recently, next-generation sequencing (NGS)-based low-coverage whole genome sequencing (LC-WGS) has been applied to detect Copy Number Gain/Loss syndromes. This dataset is intended to be used as a “Golden standard data set” for development of LC-WGS analysis method. It consists of patients (n=63) who have a mental delay and/or physical disability phenotype and normal (n=20) phenotype.
Project description:In principle, whole-genome sequencing (WGS) of the human genome even at low coverage offers higher resolution for genomic copy number variation (CNV) detection compared to array-based technologies, which is currently the first-tier approach in clinical cytogenetics. There are, however, obstacles in replacing array-based CNV detection with that of low-coverage WGS such as cost, turnaround time, and lack of systematic performance comparisons. With technological advances in WGS in terms of library preparation, instrument platforms, and data analysis algorithms, obstacles imposed by cost and turnaround time are fading. However, a systematic performance comparison between array and low-coverage WGS-based CNV detection has yet to be performed. Here, we compared the CNV detection capabilities between WGS (short-insert, 3kb-, and 5kb-mate-pair libraries) at 1X, 3X, and 5X coverages and standardly used high-resolution arrays in the genome of 1000-Genomes-Project CEU genome NA12878. CNV detection was performed using standard analysis methods, and the results were then compared to a list of Gold Standard NA12878 CNVs distilled from the 1000-Genomes Project. Overall, low-coverage WGS is able to detect drastically more (approximately 5 fold more on average) Gold Standard CNVs compared to arrays and is accompanied with fewer CNV calls without secondary validation. Furthermore, we also show that WGS (at ≥1X coverage) is able to detect all seven validated deletions larger than 100 kb in the NA12878 genome whereas only one of such deletions is detected in most arrays. Finally, we show that the much larger 15 Mbp Cri-du-chat deletion can be clearly seen at even 1X coverage from short-insert WGS.
Project description:The fallopian tube (FT) has been proposed as a potential site of origin for high-grade serous ovarian cancer (HGSOC), supporting investigation of genomic alterations across matched tissues. This dataset includes whole-genome sequencing (WGS) and DigiPico data from matched samples, including peripheral blood mononuclear cells (PBMCs), fallopian tube tissue, and tumor tissue from HGSOC patients. The data support analysis of germline and somatic variants, copy number alterations (CNAs), and neoantigen prediction across matched sample types. This submission contains the WGS data DigiPico data associated with this study.
2026-04-07 | GSE301180 | GEO
Project description:Low-coverage WGS dataset of the Tetrix japonica grasshopper complex
Project description:This dataset holds three runs of our new Atrandi-SPC based single-cell WGS protocol, one for each lysis protocol (R1-R3 in the study)
Project description:Low coverage whole genome sequencing (lc-WGS) from inducible Tet TKO (Tet iTKO) and control (Ctrl) mouse ESCs (mESC), as well as for germline Dnmt TKO mESCs. mESCs were sorted to isolate the Live/Dead dye and Thy1.2 negative CD326+GFP+ population representing the mESCs populations responsive to the tamoxifen treatment. The cells were resuspended in FACS buffer and filtered with a 70 µM filter before sorting. These bulk-population samples were analyzed by using low coverage Whole Genome Sequencing (lc-WGS).
Project description:Whole genome sequencing (WGS) of tongue cancer samples and cell line was performed to identify the fusion gene translocation breakpoint. WGS raw data was aligned to human reference genome (GRCh38.p12) using BWA-MEM (v0.7.17). The BAM files generated were further analysed using SvABA (v1.1.3) tool to identify translocation breakpoints. The translocation breakpoints were annotated using custom scripts, using the reference GENCODE GTF (v30). The fusion breakpoints identified in the SvABA analysis were additionally confirmed using MANTA tool (v1.6.0).
Project description:The current RNA sequencing dataset (n=345) provides a transcriptional map of cultured primary human fibroblasts (in vitro) aged through the cellular lifespan. The dataset was generated from paired-end (150bp) Illumina HiSeq sequencing, with an average coverage of 38.3million reads per sample. These data contain gene expression trajectories collected at high temporal frequency (every 11 days, ~7timepoints per donor) across the replicative lifespan of primary human fibroblast cell lines from n=4 healthy male and female donors, and n=3 donors with lifespan-altering mitochondrial disease caused by mutations in the SURF1 gene. The dataset can be integrated with other measures collected in parallel along the lifespan, including cytological (cell size, morphology), bioenergetic (energy expenditure, derived ATP synthesis rates from Seahorse), epigenetic (DNA methylation, GSE179847), secreted proteins, telomere length, and whole-genome sequencing (WGS) data. This dataset also includes experimental manipulations with treatments targeting oxidative phosphorylation (OXPHOS) and glycolysis, and glucocorticoid signaling, providing an opportunity to examine the influence of stress and bioenergetics on human aging biology. All processed data can be accessed and browsed at our webtool: https://columbia-picard.shinyapps.io/shinyapp-Lifespan_Study/