Unknown

Dataset Information

0

PaIRKAT: A pathway integrated regression-based kernel association test with applications to metabolomics and COPD phenotypes.


ABSTRACT: High-throughput data such as metabolomics, genomics, transcriptomics, and proteomics have become familiar data types within the "-omics" family. For this work, we focus on subsets that interact with one another and represent these "pathways" as graphs. Observed pathways often have disjoint components, i.e., nodes or sets of nodes (metabolites, etc.) not connected to any other within the pathway, which notably lessens testing power. In this paper we propose the Pathway Integrated Regression-based Kernel Association Test (PaIRKAT), a new kernel machine regression method for incorporating known pathway information into the semi-parametric kernel regression framework. This work extends previous kernel machine approaches. This paper also contributes an application of a graph kernel regularization method for overcoming disconnected pathways. By incorporating a regularized or "smoothed" graph into a score test, PaIRKAT can provide more powerful tests for associations between biological pathways and phenotypes of interest and will be helpful in identifying novel pathways for targeted clinical research. We evaluate this method through several simulation studies and an application to real metabolomics data from the COPDGene study. Our simulation studies illustrate the robustness of this method to incorrect and incomplete pathway knowledge, and the real data analysis shows meaningful improvements of testing power in pathways. PaIRKAT was developed for application to metabolomic pathway data, but the techniques are easily generalizable to other data sources with a graph-like structure.

SUBMITTER: Carpenter CM 

PROVIDER: S-EPMC8565741 | biostudies-literature | 2021 Oct

REPOSITORIES: biostudies-literature

altmetric image

Publications

PaIRKAT: A pathway integrated regression-based kernel association test with applications to metabolomics and COPD phenotypes.

Carpenter Charlie M CM   Zhang Weiming W   Gillenwater Lucas L   Severn Cameron C   Ghosh Tusharkanti T   Bowler Russell R   Kechris Katerina K   Ghosh Debashis D  

PLoS computational biology 20211022 10


High-throughput data such as metabolomics, genomics, transcriptomics, and proteomics have become familiar data types within the "-omics" family. For this work, we focus on subsets that interact with one another and represent these "pathways" as graphs. Observed pathways often have disjoint components, i.e., nodes or sets of nodes (metabolites, etc.) not connected to any other within the pathway, which notably lessens testing power. In this paper we propose the Pathway Integrated Regression-based  ...[more]

Similar Datasets

| S-EPMC10601228 | biostudies-literature
| S-EPMC4724299 | biostudies-literature
| S-EPMC2429859 | biostudies-literature
| S-EPMC4570290 | biostudies-literature
| S-EPMC4158946 | biostudies-other
| S-EPMC10174706 | biostudies-literature
| S-EPMC5860324 | biostudies-literature
2008-06-21 | E-TABM-289 | biostudies-arrayexpress
| S-EPMC3777709 | biostudies-literature
| S-EPMC10598548 | biostudies-literature