{"database":"NODE","file_versions":[],"scores":null,"additional":{"omics_type":["Genomics"],"submitter":["Lingling Li"],"technology_type":["WGS"],"disease":["['obsolete Recurrent Duodenal Cancer']"],"full_dataset_link":["https://www.biosino.org/node/experiment/detail/OEX00021328"],"experiment_platform":["Illumina NovaSeq 6000"],"experiment_library_layout":["Paired"],"experiment_library_selection":["PCR"],"sample_count":["120"],"tissue":["['duodenum']"],"taxonomy":["['Homo sapiens']"],"experiment_protocol":["DNA extraction and DNA qualification. All the samples were firstly dewaxing with dimethylbenzene, and then DNA degradation and contamination were monitored on 1% agarose gels. And DNA concentration was measured byQubit® DNA Assay in Qubit® 2.0 Flurometer (Invitrogen, USA). A total amount of at least 0.6 μg genomic DNA per sample was used as input for DNA sample preparation. Library preparation. A total amount of 0.6 μg genomic DNA per sample was used as input for DNA sample preparation. Sequencing libraries were generated using Agilent SureSelect Human All Exon kit (Agilent Technologies, CA, USA) following manufacturer’s recommendations and index codes were added to each sample. Fragmentation was carried out by hydrodynamic shearing system (Covaris, Massachusetts, USA) to generate randomly 180-280 bp fragments. Remaining overhangs were converted into blunt ends via exonuclease/polymerase activities. After adenylation of 3’ ends of DNA fragments, adapter oligonucleotides were ligated. DNA fragments with ligated adapter molecules on both ends were selectively enriched in a PCR reaction. After PCR reaction, libraries hybridize with liquid phase with biotin labeled probe, then use magnetic beads with streptomycin to capture the exons of genes. Captured libraries were enriched in a PCR reaction to add index tags to prepare for sequencing. Products were purified using AMPure XP system (Beckman Coulter, Beverly, USA) and quantified using the Agilent high sensitivity DNA assay on the Agilent Bioanalyzer 2100 system. The clustering of the index-coded samples was performed on a cBot Cluster Generation System using Hiseq PE Cluster Kit (Illumina) according to the manufacturer’s instructions. After cluster generation, the DNA libraries were sequenced on Illumina Hiseq platform and 150 bp paired-end reads were generated. Quality control of data processing and analysis. Paired-end sequencing (PE150) was performed on an Illumina HiSeq (Illumina Novaseq 6000). The resulting sequence libraries (the paired-end sequence and insert DNA between two ends) were quantified with a Qubit 2.0 (Thermo Fisher) and insert size was determined using an Agilent 2100 Bioanalyzer. The original fluorescence image files obtained from Hiseq platform are transformed to short reads (raw data) by base calling and these short reads are recorded in FASTQ format, which contains sequence information and corresponding sequencing quality information. Base calling was used to obtain the raw data (sequenced reads) from the primary image data. Quality control: (1) Discard a paired read if either one read contains adapter contamination (>10 nucleotides aligned to the adapter, allowing ≤ 10% mismatches; (2) Discard a paired read if more than 10% of bases are uncertain in either one read; (3) Discard a paired read if the proportion of low quality (Phred quality < 5) bases is over 50% in either one read. All the downstream bioinformatics analyses were based on the high-quality clean data, which were retained after these steps. At the same time, QC statistics including total reads number, raw data, raw depth, sequencing error rate, percentage of reads with Q30 (the percent of bases with phred-scaled quality score greater than 30) and QC content distribution were calculated and summarized. Reads mapping to reference sequence. Valid sequencing data was mapped to the reference human genome (UCSC hg19) by Burrows-Wheeler Aligner (BWA) software94 to get the original mapping results stored in BAM format. If one or one paired read(s) were mapped to multiple positions, the strategy adopted by BWA was to choose the most likely placement. If two or more most likely placements presented, BWA picked one randomly. Then, SAMtools95 and Picard (http://broadinstitute.github.io/picard/) were used to sort BAM files and do duplicate marking, local realignment, and base quality recalibration to generate final BAM file for computation of the sequence coverage and depth. Mapping step was very difficult due to mismatches, including true mutation and sequencing error, and duplicates resulted from PCR amplification. These duplicate reads were uninformative and shouldn’t be considered as evidence for variants. These duplicate reads were uninformative and shouldn’t be considered as evidence for variants. We used Picard to mark these duplicates for follow up analysis. Detecting and callings of somatic mutations. BWA and Samblaster were used to genome alignment, and muTect Software96 was used for targeting Somatic SNV sites, and Strelka97 was used to test Somatic INDEL information. Statistics used in the manuscript includes moderated t-statistics, and Fisher’s exact test. The manuscript statistics were used to moderated t-test and Fisher’s exact test."],"repository":["NODE"],"additional_accession":[]},"is_claimable":false,"name":"OEX_Lingling_2302201447","description":"One hundred and twenty samples of 47 duodenal cancer cases were analyzed by WES including the D/G/P/N/C subtypes. Paired-end sequencing (PE150) was performed on an Illumina HiSeq with a 135× target depth (mean) and 15.5G volume (mean) of 120 samples' raw data.","dates":{"publication":"2023-06-19","submission":"2023-02-20"},"accession":"OEX00021328","cross_references":{"NODE":["OEP00003870"]}}