Project description:Low-pass sequencing (sequencing a genome to an average depth less than 1× coverage) combined with genotype imputation has been proposed as an alternative to genotyping arrays for trait mapping and calculation of polygenic scores. To empirically assess the relative performance of these technologies for different applications, we performed low-pass sequencing (targeting coverage levels of 0.5× and 1×) and array genotyping (using the Illumina Global Screening Array (GSA)) on 120 DNA samples derived from African and European-ancestry individuals that are part of the 1000 Genomes Project. We then imputed both the sequencing data and the genotyping array data to the 1000 Genomes Phase 3 haplotype reference panel using a leave- one-out design. We evaluated overall imputation accuracy from these different assays as well as overall power for GWAS from imputed data, and computed polygenic risk scores for coronary artery disease and breast cancer using previously derived weights. We conclude that low-pass sequencing plus imputation, in addition to providing a substantial increase in statistical power for genome wide association studies, provides increased accuracy for polygenic risk prediction at effective coverages of ∼ 0.5× and higher compared to the Illumina GSA.
Project description:Purpose: Genome wide association studies (GWAS) have identified 14 loci for atrial fibrillation, but the mechanisms responsible for these associations as well as the causal genetic variants remain enigmatic. Genetic variants altering expression levels of nearby genes is one such plausible mechanism. Methods: We performed polyA+ RNA sequencing of left atrial appendages from a biracial cohort of 265 subjects. Approximately 50 million read fragments mapped to the transcriptome using TopHat aligner to hg19 and fragments counted against Ensembl 71 reference transcriptome with htseq. We also obtained genotypes using Illumina Hap550v3 and Hap610-quad SNP microarrays. Genotypes were then imputed to 1000 Genomes using IMPUTE. Expression surrogate variables were calculated using the R package sva and used as covariates in a genome-wide eQTL scan. Provision of raw data for this study, even to dbGaP, was not permitted per the study IRB. Therefore, raw data are not available for this study.