Project description:Sequence-to-function neural networks learn cis-regulatory sequence rules driving many types of genomic data. Interpreting these models to relate the sequence rules to underlying biological processes remains challenging, especially for complex genomic readouts such as MNase-seq, which maps nucleosome occupancy but is confounded by experimental bias. Here, we introduce pairwise influence by sequence attribution (PISA), which uses attribution to combinatorially decode which bases contributed to the readout at a specific genomic coordinate. PISA visualizes the effects of transcription factor motifs, detects undiscovered motifs with complex contribution patterns, and reveals experimental biases. By learning the bias for MNase-seq, PISA enables unprecedented nucleosome prediction models. These models allow the de novo discovery of nucleosome-positioning motifs and reveal the basis of Micro-C chromatin domain boundaries through systematic motif perturbations. Finally, these models allow the design of sequences with altered nucleosome configurations. These results show that PISA is a versatile tool that expands our ability to train and interpret sequence-to-function neural networks on genomics data and understand the underlying cis-regulatory code.
Project description:PISA experimental assays investigating structural alterations in Light vs Dark growth conditions with Dense and Dilute conditions. Samples are representative of Synechococcus elongatus PCC 7942 cell lysates (BTO:0004304). Purpose: Detecting structural and abundance changes in cyanobacterial cultures grown in conditions mimicking environmental perturbations such as dense light, dense dark, dilute light, and dilute dark. Technique: Proteome Integral Solubility Assay (PISA) and Global Proteomics. Other Details: TMT10 Labelled, Not-fractionated, 3 plexes in total (Plex1- PISA Dense Light and Dark, Plex2- PISA Dilute Light and Dark, Plex3- Global Dilute Light and Dark, No global proteomics for dense samples). Data was searched with MS-GF+ using PNNL's DMS Processing pipeline.
Project description:Analyses of new genomic, transcriptomic or proteomic data commonly result in trashing many unidentified data escaping the ‘canonical’ DNA-RNA-protein scheme. Testing systematic exchanges of nucleotides over long stretches produces inversed RNA pieces (here named “swinger” RNA) differing from their template DNA. These may explain some trashed data. Here analyses of genomic, transcriptomic and proteomic data of the pathogenic Tropheryma whipplei according to canonical genomic, transcriptomic and translational 'rules' resulted in trashing 58.9% of DNA, 37.7% RNA and about 85% of mass spectra (corresponding to peptides). In the trash, we found numerous DNA/RNA fragments compatible with “swinger” polymerization. Genomic sequences covered by «swinger» DNA and RNA are 3X more frequent than expected by chance and explained 12.4 and 20.8% of the rejected DNA and RNA sequences, respectively. As for peptides, several match with “swinger” RNAs, including some chimera, translated from both regular, and «swinger» transcripts, notably for ribosomal RNAs. Congruence of DNA, RNA and peptides resulting from the same swinging process suggest that systematic nucleotide exchanges increase coding potential, and may add to evolutionary diversification of bacterial populations.