Project description:It is widely recognized that the missing heritability of many human diseases is partially due to noncoding genetic variants, but there are multiple challenges that hinder the identification of functional disease-associated noncoding variants. The number of noncoding variants can be many times of coding variants; many of them are not functional but in linkage disequilibrium with the functional ones; different variants can have epistatic effects; different variants can affect the same genes or pathways in different individuals, and some variants are related to each other not by affecting the same gene but by affecting the binding of the same upstream regulator. To overcome these difficulties, we propose a novel analysis framework that considers convergent impacts of different genetic variants on protein binding, which provides multi-granular information about disease-associated perturbations of regulatory elements, genes, and pathways. Applying it to our whole-genome sequencing data of 918 short-segment Hirschsprung disease patients and matched controls, we identify various novel genes not detected by standard single-variant and region-based tests, functionally centering on neural crest migration and development. Our framework also identifies upstream regulators whose binding is influenced by the noncoding variants. Using human neural crest cells, we confirm cell-stage-specific regulatory roles three top novel regulatory elements on our list, respectively in the RET, RASGEF1A and PIK3C2B loci. In the PIK3C2B regulatory element, we further show that a noncoding variant found only in the affects the binding of the gliogenesis regulator NFIA, with a corresponding down-regulation of multiple genes in the same topologically associating domain.
Project description:Interventions: The aim of this research is to find the undiscovered gene or genes responsible for hereditary colorectal cancer, and to understand how the gene(s) cause colorectal cancer. One of the best approaches used to find the answers to these questions is to study families where there are several family members affected with colorectal cancer.
Case probands and their relatives (affected and unaffected) will be recruited
(20 ml) of blood will be obtained on one occasion from each participant
Where tissues from affected family members have already been obtained (paraffin embedded tumour tissue in blocks) we will seek permission to access this material for DNA extraction. The duration of study will be approximately 48 months.
Primary outcome(s): To identify molecular markers of potential value in understanding predisposition to CRC by their association with gene variants that increase the risk of CRC.
The study will be a family based analysis of SNP genotypes for cases and controls within families and between families. Transmission of disease associated SNPs will be assessed to determine within-family association and linkage and then to calculate across-family associations.[2-3 blood samples per week over 36 months approx]
Study Design: Purpose: Natural history;Duration: Cross-sectional;Selection: Defined population;Timing: Both
Project description:To efficiently identify genetic susceptibility variants for gastric cancer, including rare coding variants, we performed an exome chip-based array study. We found that a linkage disequilibrium (LD) block containing 2 significant variants in PSCA gene increased the risk and two blocks that included 15 suggested variants including TRIM31, TRIM 40, TRIM 10, and TRIM26 regions, and included one suggested variant and OR2H2 gene showed protective associations with gastric cancer susceptibility. In addition, the PLEC region (rs200893203), FBLN2 region (rs201192415), and EPHA2 region (rs3754334) were associated with increased susceptibility
Project description:A number of genetic studies have identified rare protein-coding DNA variations associated with autism spectrum disorder (ASD), a neurodevelopmental disorder with significant genetic etiology and heterogeneity. In contrast, the contributions of functional, regulatory genetic variations that occur in the extensive non-protein-coding regions of the genome remain poorly understood. Here we developed a genome-wide analysis to identify rare single nucleotide variants (SNVs) that occur in non-coding regions and determined regulatory function and evolutionary conservation of these variants. Using publicly available datasets and computational predictions, we identified SNVs within putative regulatory regions in promoters, transcription factor binding sites, microRNA genes and their target sites. Overall, we found regulatory variants in the ASD cases were enriched in autism-risk genes and genes involved in fetal neurodevelopment. As with previously reported coding mutations, we found an enrichment of regulatory variants associated with dysregulation of neurodevelopmental and synaptic signaling pathways. Among these were rare inherited non-coding SNVs found in the mature sequence of a number of microRNAs predicted to affect the regulation of autism-risk genes. We show a paternally inherited miR-873-5p variant, with reduced NRXN2 binding affinity, overlays a maternally inherited NRXN1 putative loss-of-function coding variation to likely increase genetic liability in an idiopathic ASD case. Our analysis pipeline provides a new resource for identifying loss-of-function regulatory DNA variations that may contribute to the genetic etiology of complex disorders.
Project description:Most DNA variants associated with common complex diseases fall outside the protein-coding regions of the genome, making them hard to detect and relate to a function. Although many computational tools are available for prioritizing functional disease risk variants outside the protein-coding regions of the genome, the precision of prediction of these tools is mostly unreliable and hence not close to cancer risk prediction. This study brings to light a novel way to improve prediction accuracy of publicly available tools by integrating the impact of cis-overlapping binding sites of opposing cancer proteins, such as P53 and cMYC, in their analysis to filter out deleterious DNA variants outside the protein-coding regions of the human genome. Using a biology-based statistical approach, DNA variants within cis-overlapping motifs impacting the binding affinity of opposing transcription factors can significantly alter the expression of target genes and regulatory networks. This study brings us closer to developing a generally applicable approach capable of filtering etiological non-coding variations in co-occupied genomic regions of P53 and cMYC family members to improve disease risk assessment.
Project description:To efficiently identify genetic susceptibility variants for gastric cancer, including rare coding variants, we performed an exome chip-based array study. We found that a linkage disequilibrium (LD) block containing 2 significant variants in PSCA gene increased the risk and two blocks that included 15 suggested variants including TRIM31, TRIM 40, TRIM 10, and TRIM26 regions, and included one suggested variant and OR2H2 gene showed protective associations with gastric cancer susceptibility. In addition, the PLEC region (rs200893203), FBLN2 region (rs201192415), and EPHA2 region (rs3754334) were associated with increased susceptibility We performed an exome chip-based array study in 329 gastric cancer cases and 683 controls.
Project description:This data set includes the following summary level data files used for the 13k analysis of T2D-GENES data: wes.variants.list: list of variants to keep for any analysis of the exomes data wes.assoc.samples.list: list of samples to keep for association analysis wes.assoc.variants.list: list of variants to keep for association analysis wes.sv.assoc.txt: single variant association analysis results wes.gene.ptv.variants.list.txt: list of protein truncating variants to use in gene-level analysis wes.gene.ptv.assoc.txt: results from gene-level tests of protein truncating variants wes.gene.nsstrict.variants.list.txt: list of NSstrict variants to use in gene-level analysis wes.gene.nsstrict.assoc.txt: results from gene-level tests of NSstrict variants wes.gene.nsbroad.variants.list.txt: list of NSbroad variants to use in gene-level analysis wes.gene.nsbroad.assoc.txt: results from gene-level tests of NSbroad variants wes.gene.ns.variants.list.txt: list of non synonymous variants to use in gene-level analysis wes.gene.ns.assoc.txt: results from gene-level tests of non synonymous variants
Project description:Homeodomains (HDs) are the second largest class of DNA binding domains (DBDs) in eukaryotic sequence-specific transcription factors (TFs) and are the TF structural class with the largest number of disease mutations in the Human Gene Mutation Database (HGMD). Despite numerous structural studies and large-scale analyses of HD DNA binding specificity, HD-DNA recognition is still not fully understood. Here, we analyzed 92 human HD mutants, including disease-associated variants and variants of unknown significance (VUS), for their effects on DNA binding activity. Many of the variants altered DNA binding affinity and/or specificity. Structural analysis identified 14 novel specificity-determining positions, 5 of which do not contact DNA. The same missense substitution at analogous positions within different HDs exhibited different effects on DNA binding activity. Variant effect prediction tools perform moderately well in distinguishing variants with altered DNA binding affinity, but poorly in identifying those with altered binding specificity. Our results highlight the need for biochemical assays of TF coding variants and promote dozens of variants for further investigations into their pathogenicity and the development of clinical diagnostics and precision therapies.
Project description:Interventions: Group 1: Blood and sputum samples as well as paraffin embedded tumour tissue from patients with microsatellite stable colorectal cancer shall be analysed before therapy and over time to establish and validate hotspot mutation and somatic copy number variant (SCNAs) analysis. We therefore need the following samples:
- Two Cell-Free DNA BCT CE Streck tubes with 8 ml blood per sampling for preparation of plasma-DNA.
- One sputum tube (only at study inclusion).
- FFPE tissue samples from the primary tumor (from initial surgery)
Group 2: Blood samples of tumor-free control persons shall be tested for hotspot mutations and somatic copy number variants (SCNAs) to identify technical artefacts and improve our protocols. We therefore need the following samples:
- Two Cell-Free DNA BCT CE Streck tubes with 8 ml blood per sampling for preparation of plasma-DNA.
Primary outcome(s): Identification of tumorspecific SCNAs in plasma samples of colorectal cancer patients
Study Design: Allocation: ; Masking: ; Control: ; Assignment: ; Study design purpose: other
Project description:Specific DNA-protein interactions mediate physiologic gene regulation and may be altered by DNA variants linked to polygenic disease. To enhance the speed and signal-to-noise ratio (SNR) of identifying and quantifying proteins associating with specific DNA sequences in living cells, we developed proximal biotinylation by episomal recruitment (PROBER). PROBER uses high copy episomes to amplify SNR along with proximity proteomics (BioID) to identify the transcription factors (TFs) and additional gene regulators associated with DNA sequences of interest. PROBER quantified steady-state and inducible association of TFs and associated chromatin regulators to target DNA sequences and quantified binding quantitative trait loci (bQTLs) due to single nucleotide variants. PROBER identified alterations in gene regulator associations due to cancer hotspot mutations in the hTERT promoter, indicating these mutations increase promoter association with specific gene activators. PROBER offers an approach to rapidly identify proteins associated with specific DNA sequences and their variants in living cells.