Project description:Whole exome sequencing of 5 HCLc tumor-germline pairs. Genomic DNA from HCLc tumor cells and T-cells for germline was used. Whole exome enrichment was performed with either Agilent SureSelect (50Mb, samples S3G/T, S5G/T, S9G/T) or Roche Nimblegen (44.1Mb, samples S4G/T and S6G/T). The resulting exome libraries were sequenced on the Illumina HiSeq platform with paired-end 100bp reads to an average depth of 120-134x. Bam files were generated using NovoalignMPI (v3.0) to align the raw fastq files to the reference genome sequence (hg19) and picard tools (v1.34) to flag duplicate reads (optical or pcr), unmapped reads, reads mapping to more than one location, and reads failing vendor QC.
Project description:The expression profile and sequence variants of 476 early stage urothelial carcinoma were studied using whole transcriptome sequencing. RNA-Seq libraries were prepared by Ribo-Zero treatment of total-RNA (to reduce the rRNA content) followed by library preparation using ScriptSeq. RNA-Seq libraries were paired-end sequenced (2x 101 bp) on Illumina HiSeq 2000 and the resulting fastq files were processed using tools from the Genome Analysis Toolkit (GATK and from the Tuxedo suite. Access to the sequence data (bam and vcf files), containing person identifying information, needs signature on a controlled access form, and can be accessed at The European Genome-phenome Archive (EGA) using the study ID EGAS00001001236 following request. An expression matrix of FPKM values are available without restriction at ArrayExpress.
Project description:CTCF ChIP-seq of 39 primary samples derived from human acute leukemias, namely AML, T-ALL and mixed myeloid/lymphoid leukemias with CpG Island Methylator Phenotype (CIMP). Due to patient confidentiality considerations, the raw data files for this dataset have been deposited to the EGA controlled-access archive under the accession numbers EGAS00001007094 (study); EGAD00001011059 (dataset).
Project description:Whole genome sequencing (WGS) of tongue cancer samples and cell line was performed to identify the fusion gene translocation breakpoint. WGS raw data was aligned to human reference genome (GRCh38.p12) using BWA-MEM (v0.7.17). The BAM files generated were further analysed using SvABA (v1.1.3) tool to identify translocation breakpoints. The translocation breakpoints were annotated using custom scripts, using the reference GENCODE GTF (v30). The fusion breakpoints identified in the SvABA analysis were additionally confirmed using MANTA tool (v1.6.0).
Project description:By generating a paired single cell RNA-sequencing database of the tumor niche from 10 newly diagnosed MM patients, we created a unique dataset allowing the in-depth analyses of stromal-immune interactions within the tumor microenvironment (see related accession number). Using this database, we identified the presence of inflammatory stromal fibroblasts in the bone marrow of Myeloma patients.The stromal inflammation was associated with NF-κB signaling, and sources of IL-1β or TNFα were specific immune subsets previously shown to be altered in MM, suggesting the presence of an immune cell-mediated feed-forward loop of bone marrow inflammation in MM. By tracking inflammatory signatures over time in individual patients undergoing first-line treatment using bulk RNA sequencing, we show that bone marrow inflammation is not reverted by successful anti-tumor therapy (this dataset), suggesting a role for stromal fibroblasts and bone marrow inflammation in disease persistence or relapse. Raw sequencing data files will be deposited to EGA.
Project description:H3K27ac ChIP-seq of 79 primary samples derived from human acute leukemias, namely AML, T-ALL and mixed myeloid/lymphoid leukemias with CpG Island Methylator Phenotype (CIMP). In addition, 4 samples derived from CD34+ cord blood cells of healthy donors were included. Due to patient confidentiality considerations, the raw data files for this dataset have been deposited to the EGA controlled-access archive under the accession numbers EGAS00001007094 (study); EGAD00001011060 (dataset).
Project description:Hi-C of 17 primary samples obtained from human acute leukemias, namely AML, T-ALL and mixed myeloid/lymphoid leukemias with CpG Island Methylator Phenotype (CIMP). As healthy controls, Hi-C of CD34+ HSPCs from 3 healthy donors were used. Due to patient confidentiality considerations, the raw data files for this dataset have been deposited to the EGA controlled-access archive under the accession numbers EGAS00001007094 (study); EGAD00001011051 (dataset).
Project description:Cancer cell lines can provide robust and facile biological models for the generation and testing of hypothesis in the early stages of drug development and caner biology. Although clinical trials remain the ultimate scientific testing ground for anticancer therapies, the use of appropriate model systems to explore the molecular basis of drug activity and to identify predictive biomarkers during their development can have a profound effect on the design, cost and ultimate success of new cancer drug development. In order to capture the high degree of genomic diversity in cancer and to identify rare molecular subtypes, we have assembled a collection of >1000 cancer cell lines. These lines have been characterised using whole exome sequencing, genome wide analysis of copy number, mRNA gene expression profiling and DNA methylation analysis (http://cancer.sanger.ac.uk/cell_lines). To further characterise this panel of cell lines we have now compiled data for RNA sequencing. The current study represent data for ~450 of the cell lines in the panel, data for the remaining lines can be accessed via the CGHUB data browser hosted at UCSC. <br>This ArrayExpress record contains only meta-data. Raw data files have been archived at the European Genome-Phenome Archive (EGA, www.ebi.ac.uk/ega) by the consortium, with restricted access to protect sample donors' identity. The relevant accessions of the EGA data set is EGAD00001001357 under EGA study accession EGAS00001000828.