Evaluation of variable selection methods for random forests and omics data sets.
Ontology highlight
ABSTRACT: Machine learning methods and in particular random forests are promising approaches for prediction based on high dimensional omics data sets. They provide variable importance measures to rank predictors according to their predictive power. If building a prediction model is the main goal of a study, often a minimal set of variables with good prediction performance is selected. However, if the objective is the identification of involved variables to find active networks and pathways, approaches that aim to select all relevant variables should be preferred. We evaluated several variable selection procedures based on simulated data as well as publicly available experimental methylation and gene expression data. Our comparison included the Boruta algorithm, the Vita method, recurrent relative va
SUBMITTER: Degenhardt F
PROVIDER: S-EPMC6433899 | biostudies-literature | 2019 Mar
REPOSITORIES: biostudies-literature
ACCESS DATA