{"database":"biostudies-literature","file_versions":[],"scores":null,"additional":{"submitter":["Srivastava H"],"funding":["NIH Office of the Director","National Heart, Lung, and Blood Institute","NHLBI NIH HHS","NIH HHS"],"pagination":["e1010702"],"full_dataset_link":["https://www.ebi.ac.uk/biostudies/studies/S-EPMC9681107"],"repository":["biostudies-literature"],"omics_type":["Unknown"],"volume":["18(11)"],"pubmed_abstract":["Protein and mRNA levels correlate only moderately. The availability of proteogenomics data sets with protein and transcript measurements from matching samples is providing new opportunities to assess the degree to which protein levels in a system can be predicted from mRNA information. Here we examined the contributions of input features in protein abundance prediction models. Using large proteogenomics data from 8 cancer types within the Clinical Proteomic Tumor Analysis Consortium (CPTAC) data set, we trained models to predict the abundance of over 13,000 proteins using matching transcriptome data from up to 958 tumor or normal adjacent tissue samples each, and compared predictive performances across algorithms, data set sizes, and input features. Over one-third of proteins (4,648) showe"],"journal":["PLoS computational biology"],"pubmed_title":["Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners."],"pmcid":["PMC9681107"],"funding_grant_id":["R00 HL127302","R01 HL141278","R00-HL127302","R00-HL144829","R03 OD032666","R00 HL144829","R01-HL141278","R03-OD032666;"],"pubmed_authors":["Lippincott MJ","Lam MPY","Srivastava H","Lau E","Canfield R","Currie J"],"additional_accession":[]},"is_claimable":false,"name":"Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners.","description":"Protein and mRNA levels correlate only moderately. The availability of proteogenomics data sets with protein and transcript measurements from matching samples is providing new opportunities to assess the degree to which protein levels in a system can be predicted from mRNA information. Here we examined the contributions of input features in protein abundance prediction models. Using large proteogenomics data from 8 cancer types within the Clinical Proteomic Tumor Analysis Consortium (CPTAC) data set, we trained models to predict the abundance of over 13,000 proteins using matching transcriptome data from up to 958 tumor or normal adjacent tissue samples each, and compared predictive performances across algorithms, data set sizes, and input features. Over one-third of proteins (4,648) showe","dates":{"release":"2022-01-01T00:00:00Z","publication":"2022 Nov","modification":"2025-04-03T23:17:11.855Z","creation":"2025-04-03T23:17:11.855Z"},"accession":"S-EPMC9681107","cross_references":{"pubmed":["36356032"],"doi":["10.1371/journal.pcbi.1010702"]}}