Supervised dimensionality reduction for big data.
Ontology highlight
ABSTRACT: To solve key biomedical problems, experimentalists now routinely measure millions or billions of features (dimensions) per sample, with the hope that data science techniques will be able to build accurate data-driven inferences. Because sample sizes are typically orders of magnitude smaller than the dimensionality of these data, valid inferences require finding a low-dimensional representation that preserves the discriminating information (e.g., whether the individual suffers from a particular disease). There is a lack of interpretable supervised dimensionality reduction methods that scale to millions of dimensions with strong statistical theoretical guarantees. We introduce an approach to extending principal components analysis by incorporating class-conditional moment estimates into the
SUBMITTER: Vogelstein JT
PROVIDER: S-EPMC8129083 | biostudies-literature | 2021 May
REPOSITORIES: biostudies-literature
ACCESS DATA