Efficient nonparametric statistical inference on population feature importance using Shapley values.
Ontology highlight
ABSTRACT: The true population-level importance of a variable in a prediction task provides useful knowledge about the underlying data-generating mechanism and can help in deciding which measurements to collect in subsequent experiments. Valid statistical inference on this importance is a key component in understanding the population of interest. We present a computationally efficient procedure for estimating and obtaining valid statistical inference on the Shapley Population Variable Importance Measure (SPVIM). Although the computational complexity of the true SPVIM scales exponentially with the number of variables, we propose an estimator based on randomly sampling only Θ(n) feature subsets given n observations. We prove that our estimator converges
SUBMITTER: Williamson BD
PROVIDER: S-EPMC8057672 | biostudies-literature | 2020 Jul
REPOSITORIES: biostudies-literature
ACCESS DATA