Optimizing differential expression analysis for proteomics data via high-performing rules and ensemble inference.
Ontology highlight
ABSTRACT: Identification of differentially expressed proteins in a proteomics workflow typically encompasses five key steps: raw data quantification, expression matrix construction, matrix normalization, missing value imputation (MVI), and differential expression analysis. The plethora of options in each step makes it challenging to identify optimal workflows that maximize the identification of differentially expressed proteins. To identify optimal workflows and their common properties, we conduct an extensive study involving 34,576 combinatoric experiments on 24 gold standard spike-in datasets. Applying frequent pattern mining techniques to top-ranked workflows, we uncover high-performing rules that demonstrate optimality has conserved properties. Via machine learning, we confirm optimal workflows
SUBMITTER: Peng H
PROVIDER: S-EPMC11082229 | biostudies-literature | 2024 May
REPOSITORIES: biostudies-literature
ACCESS DATA