Unknown

Dataset Information

0

Global functional atlas of Escherichia coli encompassing previously uncharacterized proteins.


ABSTRACT: One-third of the 4,225 protein-coding genes of Escherichia coli K-12 remain functionally unannotated (orphans). Many map to distant clades such as Archaea, suggesting involvement in basic prokaryotic traits, whereas others appear restricted to E. coli, including pathogenic strains. To elucidate the orphans' biological roles, we performed an extensive proteomic survey using affinity-tagged E. coli strains and generated comprehensive genomic context inferences to derive a high-confidence compendium for virtually the entire proteome consisting of 5,993 putative physical interactions and 74,776 putative functional associations, most of which are novel. Clustering of the respective probabilistic networks revealed putative orphan membership in discrete multiprotein complexes and functional modules together with annotated gene products, whereas a machine-learning strategy based on network integration implicated the orphans in specific biological processes. We provide additional experimental evidence supporting orphan participation in protein synthesis, amino acid metabolism, biofilm formation, motility, and assembly of the bacterial cell envelope. This resource provides a "systems-wide" functional blueprint of a model microbe, with insights into the biological and evolutionary significance of previously uncharacterized proteins.

SUBMITTER: Hu P 

PROVIDER: S-EPMC2672614 | biostudies-literature | 2009 Apr

REPOSITORIES: biostudies-literature

altmetric image

Publications


One-third of the 4,225 protein-coding genes of Escherichia coli K-12 remain functionally unannotated (orphans). Many map to distant clades such as Archaea, suggesting involvement in basic prokaryotic traits, whereas others appear restricted to E. coli, including pathogenic strains. To elucidate the orphans' biological roles, we performed an extensive proteomic survey using affinity-tagged E. coli strains and generated comprehensive genomic context inferences to derive a high-confidence compendiu  ...[more]

Similar Datasets

| S-EPMC3019733 | biostudies-literature
| S-EPMC6237786 | biostudies-literature
2018-08-08 | GSE111095 | GEO
| S-EPMC8464067 | biostudies-literature
| S-EPMC283594 | biostudies-other
2021-08-26 | GSE159658 | GEO