Project description:The analytes qualified as biomarkers are potent tools to diagnose various diseases, monitor therapy responses, and design therapeutic interventions. The early assessment of the diverseness of human disease is essential for the speedy and cost-efficient implementation of personalized medicine. We developed g3mclass, the Gaussian mixture modeling software for molecular assay data classification. This software automates the validated multiclass classifier applicable to single analyte tests and multiplexing assays. The g3mclass achieves automation using the original semi-constrained expectation-maximization (EM) algorithm that allows inference from the test, control, and query data that human experts cannot interpret. In this study, we used real-world clinical data and gene expression datasets (ERBB2, ESR1, PGR) to provide examples of how g3mclass may help overcome the problems of over-/underdiagnosis and equivocal results in diagnostic tests for breast cancer. We showed the g3mclass output’s accuracy, robustness, scalability, and interpretability. The user-friendly interface and free dissemination of this multi-platform software aim to ease its use by research laboratories, biomedical pharma, companion diagnostic developers, and healthcare regulators. Furthermore, the g3mclass automatic extracting information through probabilistic modeling is adaptable for blending with machine learning and artificial intelligence.
Project description:There are many subtypes of dementia, and identification of diagnostic biomarkers that are minimally-invasive, low-cost, and efficient is desired. Circulating microRNAs (miRNAs) have recently gained attention as easily accessible and non-invasive biomarkers. We conducted a comprehensive miRNA expression analysis of serum samples from 1348 Japanese dementia patients, composed of four subtypes—Alzheimer’s disease (AD), vascular dementia, dementia with Lewy bodies (DLB), and normal pressure hydrocephalus—and 246 control subjects. We used this data to construct dementia subtype prediction models based on penalized regression models with the multiclass classification. We constructed a final prediction model using 46 miRNAs, which classified dementia patients from an independent validation set into four subtypes of dementia. Network analysis of miRNA target genes revealed important hub genes, SRC and CHD3, associated with the AD pathogenesis. Moreover, MCU and CASP3, which are known to be associated with DLB pathogenesis, were identified from our DLB-specific target genes. Our study demonstrates the potential of blood-based biomarkers for use in dementia-subtype prediction models. We believe that further investigation using larger sample sizes will contribute to the accurate classification of subtypes of dementia.
Project description:The optimal treatment of patients with cancer depends on establishing accurate diagnoses by using a complex combination of clinical and histopathological data. In some instances, this task is difficult or impossible because of atypical clinical presentation or histopathology. To determine whether the diagnosis of multiple common adult malignancies could be achieved purely by molecular classification, we subjected 218 tumor samples, spanning 14 common tumor types, and 90 normal tissue samples to oligonucleotide microarray gene expression analysis. The expression levels of 16,063 genes and expressed sequence tags were used to evaluate the accuracy of a multiclass classifier based on a support vector machine algorithm. Overall classification accuracy was 78%, far exceeding the accuracy of random classification (9%). Poorly differentiated cancers resulted in low-confidence predictions and could not be accurately classified according to their tissue of origin, indicating that they are molecularly distinct entities with dramatically different gene expression patterns compared with their well differentiated counterparts. Taken together, these results demonstrate the feasibility of accurate, multiclass molecular cancer classification and suggest a strategy for future clinical implementation of molecular cancer diagnostics.
Project description:Diffuse Large B Cell Lymphoma (DLBCL) is the most common lymphoid malignancy in adults. Despite being considered a single disease, DLBCL presents with variable backgrounds in terms of morphology, genetics, and biological behavior, which results in heterogeneous outcomes among patients. Although new tools have been developed for the classification and management of patients, 40% of them still have primary refractory disease or relapse. In addition, multiple factors regarding the pathogenesis of this disease remain unclear and identification of novel biomarkers is needed. In this context, recent investigations point to microRNAs as useful biomarkers in cancer as well as important players in the development of the disease. However, regarding DLBCL, up to date, there is inconsistency in the data reported. Therefore, in this work, the main goals were to determine a microRNA set with utility as biomarkers for DLBCL diagnosis, classification, prognosis and treatment response. To achieve these goals, we analyzed microRNA expression in a cohort of 78 DLBCL samples at diagnosis and 17 controls using small RNA sequencing. This way, we were able to define new microRNA expression signatures for diagnosis, classification, treatment response and prognosis. In summary, our study remarks that microRNAs could play an important role as biomarkers in diagnosis, classification, treatment response and prognosis in DLBCL.
Project description:Dementia subtype prediction models constructed by penalized regression methods for multiclass classification using serum microRNA expression data
Project description:Cerebrospinal fluid (CSF) liquid biopsies serve as a rich source of tumor-derived cell-free DNA (cfDNA) for evaluating patients with central nervous system (CNS) tumors. However, challenges stemming from trace cfDNA yields and low mutational burden have hindered sensitivity, whereas first-generation clinical assays have relied on genetic alterations as biomarkers. Leveraging the diagnostic utility of DNA methylation classification in CNS tumors, we developed M-PACT (Methylation-based Predictive Algorithm for CNS Tumors), a robust deep neural network that accurately classifies tumors from sub-nanogram input cfDNA methylomes acquired through enzymatic methylation sequencing. In addition to tumor classification, this workflow enables methylation-based cellular deconvolution and sensitive copy number variation (CNV) detection. We benchmark our methodology in pediatric CNS embryonal tumors and further demonstrate accurate classification of intra-operative CSF, balanced tumor genomes, and secondary malignancies. Altogether, we provide a blueprint for CNS tumor classification from low input cfDNA methylomes, motivating prospective validation for future clinical implementation.
Project description:The advent of large-scale single-cell chromatin accessibility profiling has accelerated our ability to map gene regulatory landscapes, but has outpaced the development of scalable software to rapidly extract biological meaning from these data. Here we present a software suite for single-cell analysis of regulatory chromatin in R (ArchR; www.ArchRProject.com) that enables fast and comprehensive analysis of single-cell chromatin accessibility data. ArchR provides an intuitive, user-focused interface for complex single-cell analyses including doublet removal, single-cell clustering and cell type identification, unified peak set generation, cellular trajectory identification, DNA element to gene linkage, transcription factor footprinting, mRNA expression level prediction from chromatin accessibility, and multi-omic integration with scRNA-seq. Enabling the analysis of over 1.2 million single cells within 8 hours on a standard Unix laptop, ArchR is a comprehensive analytical suite for end-to-end analysis of single-cell chromatin accessibility data that will accelerate the understanding of gene regulation at the resolution of individual cells.