<HashMap><database>EGA</database><scores/><additional><omics_type>Genomics</omics_type><dataset_type>N/A</dataset_type><full_dataset_link>https://ega-archive.org/datasets/EGAD00001011363</full_dataset_link><sample_count>80</sample_count><description>EGA dataset EGAD00001011363</description><repository>EGA</repository><title>Low-coverage whole genome sequencing for a highly selective cohort of severe COVID-19 patients</title></additional><is_claimable>false</is_claimable><name>83859d7e-2be4-45af-8f43-d8be586e0945 - samples</name><description>We generated a dataset consisting of 79 VCF files, and respective FASTQ and CRAM files, methodically generated using the GLIMPSE1 imputation algorithm leveraging the 1000 Genomes Project Phase 3 dataset as the reference panel of haplotypes. In total this dataset is composed of approximately 325 GB of FASTQ data, 156 GB of CRAM data, and 6 GB of VCF data. Our samples were specifically derived from sequenced DNA from a highly selective cohort of patients, mostly comprised of Iberian Populations in Spain (IBS) individuals but also containing some individuals with other genetic backgrounds, who presented severe COVID-19 symptoms during the initial wave of the SARS-CoV-2 pandemic in Madrid, Spain. On average, each VCF file in this rich dataset contains 9.49 million high-confidence single nucleotide variants [95%CI: 9.37 million - 9.61 million].</description><dates><updated>2023-10-18 15:24:38</updated></dates><accession>EGAD00001011363</accession><cross_references><TAXONOMY>9606</TAXONOMY><EGA>EGAC00001003435</EGA><EGA>EGAS00001007573</EGA></cross_references></HashMap>