<HashMap><database>biostudies-literature</database><scores/><additional><submitter>Enright JM</submitter><funding>Natural Sciences and Engineering Research Counsel</funding><pagination>msad084</pagination><full_dataset_link>https://www.ebi.ac.uk/biostudies/studies/S-EPMC10124876</full_dataset_link><repository>biostudies-literature</repository><omics_type>Unknown</omics_type><volume>40(4)</volume><pubmed_abstract>Low complexity sequences (LCRs) are well known within coding as well as non-coding sequences. A low complexity region within a protein must be encoded by the underlying DNA sequence. Here, we examine the relationship between the entropy of the protein sequence and that of the DNA sequence which encodes it. We show that they are poorly correlated whether starting with a low complexity region within the protein and comparing it to the corresponding sequence in the DNA or by finding a low complexity region within coding DNA and comparing it to the corresponding sequence in the protein. We show this is the case within the proteomes of five model organisms: Homo sapiens, Saccharomyces cerevisiae, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana. We also report a signifi</pubmed_abstract><journal>Molecular biology and evolution</journal><pubmed_title>Low Complexity Regions in Proteins and DNA are Poorly Correlated.</pubmed_title><pmcid>PMC10124876</pmcid><funding_grant_id>PGSD3-547476-2020</funding_grant_id><funding_grant_id>USRA-526761</funding_grant_id><funding_grant_id>RGPIN-202-05733</funding_grant_id><pubmed_authors>Dickson ZW</pubmed_authors><pubmed_authors>Enright JM</pubmed_authors><pubmed_authors>Golding GB</pubmed_authors></additional><is_claimable>false</is_claimable><name>Low Complexity Regions in Proteins and DNA are Poorly Correlated.</name><description>Low complexity sequences (LCRs) are well known within coding as well as non-coding sequences. A low complexity region within a protein must be encoded by the underlying DNA sequence. Here, we examine the relationship between the entropy of the protein sequence and that of the DNA sequence which encodes it. We show that they are poorly correlated whether starting with a low complexity region within the protein and comparing it to the corresponding sequence in the DNA or by finding a low complexity region within coding DNA and comparing it to the corresponding sequence in the protein. We show this is the case within the proteomes of five model organisms: Homo sapiens, Saccharomyces cerevisiae, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana. We also report a signifi</description><dates><release>2023-01-01T00:00:00Z</release><publication>2023 Apr</publication><modification>2025-04-19T07:26:07.3Z</modification><creation>2025-02-19T04:16:32.441Z</creation></dates><accession>S-EPMC10124876</accession><cross_references><pubmed>37036379</pubmed><doi>10.1093/molbev/msad084</doi></cross_references></HashMap>