Project description:Transcription factors read the genome, fundamentally connecting DNA sequence to gene expression across diverse cell types. Determining how, where, and when TFs bind chromatin will advance our understanding of gene regulatory networks and cellular behavior. The 2017 ENCODE-DREAM in vivo Transcription-Factor Binding Site (TFBS) Prediction Challenge highlighted the value of chromatin accessibility data to TFBS prediction, establishing state-of-the-art methods. Yet, while Assay-for-Transposase-Accessible-Chromatin (ATAC)-seq datasets grow exponentially, suboptimal motif scanning is commonly used for TFBS prediction from ATAC-seq. Here, we present “maxATAC”, a suite of user-friendly, deep neural network models for genome-wide TFBS prediction from ATAC-seq in any cell type. With models available for 127 human TFs, maxATAC is the largest collection of state-of-the-art TFBS models to date. maxATAC performance extends to primary cells and single-cell ATAC-seq, enabling state-of-the-art TFBS prediction in vivo. We demonstrate maxATAC’s capabilities by identifying TFBS associated with allele-dependent chromatin accessibility at atopic dermatitis genetic risk loci.
Project description:We developed an unbiased strategy for MOA prediction, called Perturbation-Specific Transcriptional Mapping (PerSpecTM), in which large-throughput expression profiling of wildtype or hypomorphic mutants, depleted for essential targets, enables a computational strategy to address this challenge. We applied PerSpecTM to perform reference-based MOA prediction based on the principle that similar perturbations, whether small molecule or genetic, will elicit similar transcriptional responses. Using this approach, we elucidated the MOAs of three new molecules with activity against Pseudomonas aeruginosa by mapping their expression profiles to those of a reference set of antimicrobial compounds with known MOAs. We also show that transcriptional responses to small molecule inhibition maps to those resulting from genetic depletion of essential targets by CRISPRi by PerSpecTM, demonstrating proof-of-concept that correlations between expression profiles of small molecule and genetic perturbations can facilitate MOA prediction when no chemical entities exist to serve as a reference. Empowered by PerSpecTM, this work lays the foundation for an unbiased, readily scalable, systematic reference-based strategy for MOA elucidation that could transform antibiotic discovery efforts.
Project description:In this work, a total of eight proteins have been identified (six specific to the human proteome and two specific to the soybean proteome) that are supported by literature to be involved in human health, specifically related to immunological and neurological pathways. Our approach involved the use of the Protein-protein Interaction Prediction Engine (PIPE4) algorithm which was specifically developed for complex inter- and cross-species prediction schemas to generate the comprehensive interactome between H. sapiens and G. max. A literature-curated list of human proteins known to be associated with the Human Allergy Response and a second literature-curated list of soybean proteins known to be associated with Soybean Allergens were used to identify candidate proteins whose interactions may be consequential to human health. This study, beyond generating the most comprehensive human-soybean interactome to date, elucidated a soybean seed interactome and identified several proteins putatively consequential to human health.