Project description:Deconvolution models are a powerful tool for extracting cell type-specific information from bulk gene expression profiles. Current methods leverage advanced machine learning models and high-resolution sequencing, like single-cell RNA-sequencing (scRNA-seq), showing promising results across diverese tissues and conditions. However, they still present important limitations: Many depend on selecting a robust reference, which can strongly affect the deconvolution. Secondly, pseudobulk data used for training and real bulk RNA-seq samples often exhibit strong distribution shifts, which are currently unaccounted for. Finally, most deconvolution approaches behave as black boxes, which can compromise the reliability of the results. Here, we present Sweetwater, an adaptive and interpretable autoencoder that efficiently deconvolves bulk samples leveraging multiple classes of reference data. Moreover, we propose an improved way of generating training data from a mixture of FACS-sorted FASTQ files, reducing platform-specific biases and outperforming current single-cell-based references. Furthermore, we introduce a gold standard dataset to facilitate fair and accurate evaluation of deconvolution approaches. Finally, we demonstrate that Sweetwater adapts effectively to deconvolved samples during training, uncovering biologically meaningful patterns and enhancing result's reliability. Sweetwater is available at https://github.com/ML4BM-Lab/Sweetwater, and we anticipate it will expedite the accurate examination of high-throughput clinical data across diverse applications.
Project description:Numerous multi-omic investigations of cancer tissue have documented varying and poor pairwise transcript:protein quantitative correlations and most deconvolution tools aiming to predict cell type proportions (cell admixture) have been developed and credentialed using transcript-level data alone. To estimate cell admixture using protein abundance data, we analyzed proteome (and transcriptome data) generated from contrived admixtures of tumor, stroma, and immune cell models or those selectively harvested from the tissue microenvironment by laser microdissection from high grade serous ovarian cancer (HGSOC) tumors. Co-quantified transcripts and proteins performed similarly to estimate stroma and immune cell admixture in two commonly used deconvolution algorithms ESTIMATE and ConsensusTME (r ≥ 0.63). Here we have developed and optimized protein-based signatures to estimate cell admixture proportions and benchmarked these using bulk tumor proteomics data from over 150 HGSOC patients. The optimized protein signatures supporting cell type proportion estimates from bulk tissue proteomic data are available at https://lmdomics.org/ProteoMixture/.
Project description:Plasma cell-free DNA (cfDNA) is a noninvasive biomarker for cell death of all organs. Deciphering the tissue origin of cfDNA can reveal abnormal cell death because of diseases, which has great clinical potential in disease detection and monitoring. Despite the great promise, the sensitive and accurate quantification of tissue-derived cfDNA remains challenging to existing methods due to the limited characterization of tissue methylation and the reliance on unsupervised methods. To fully exploit the clinical potential of tissue-derived cfDNA, here we present one of the largest comprehensive and high-resolution methylation atlas based on 521 noncancer tissue samples spanning 29 major types of human tissues. We systematically identified fragment-level tissue-specific methylation patterns and extensively validated them in orthogonal datasets. Based on the rich tissue methylation atlas, we develop the first supervised tissue deconvolution approach, a deep-learning-powered model, cfSort, for sensitive and accurate tissue deconvolution in cfDNA. On the benchmarking data, cfSort showed superior sensitivity and accuracy compared to the existing methods. We further demonstrated the clinical utilities of cfSort with two potential applications: aiding disease diagnosis and monitoring treatment side effects. The tissue-derived cfDNA fraction estimated from cfSort reflected the clinical outcomes of the patients. In summary, the tissue methylation atlas and cfSort enhanced the performance of tissue deconvolution in cfDNA, thus facilitating cfDNA-based disease detection and longitudinal treatment monitoring.
Project description:Background: Transcriptomic data from diverse measurement technologies are widely used to study tissue heterogeneity. Cell-type deconvolution, which resolves mixed transcriptomic signals into cellular components, is a key analytical approach. However, achieving accurate deconvolution across platforms remains challenging due to platform-specific experimental and technological biases. Results: We systematically benchmarked deconvolution performance using real-world cross-platform datasets and simulated data modeling distinct technological features. SpatialDecon and cell2location demonstrated the most reliable and consistent performance across both simulated and experimental settings across a broad range of technological biases. Conclusions: Our results highlight how the different deconvolution tools are affected by data properties that depend on technological differences between transcriptomic platforms. Moreover, we provide practical guidelines for selecting computational methods dependent on experimental design for robust deconvolution of cross-platform transcriptomic data.
Project description:Background: Transcriptomic data from diverse measurement technologies are widely used to study tissue heterogeneity. Cell-type deconvolution, which resolves mixed transcriptomic signals into cellular components, is a key analytical approach. However, achieving accurate deconvolution across platforms remains challenging due to platform-specific experimental and technological biases. Results: We systematically benchmarked deconvolution performance using real-world cross-platform datasets and simulated data modeling distinct technological features. SpatialDecon and cell2location demonstrated the most reliable and consistent performance across both simulated and experimental settings across a broad range of technological biases. Conclusions: Our results highlight how the different deconvolution tools are affected by data properties that depend on technological differences between transcriptomic platforms. Moreover, we provide practical guidelines for selecting computational methods dependent on experimental design for robust deconvolution of cross-platform transcriptomic data.
Project description:Bulk deconvolution with single-cell/nucleus RNA-seq data is critical for understanding heterogeneity in complex biological samples, yet the technological discrepancy across sequencing platforms limits deconvolution accuracy. To address this, we introduce an experimental design to match inter-platform biological signals, hence revealing the technological discrepancy, and then develop a deconvolution framework called DeMixSC using the better-matched, i.e., benchmark, data. Built upon a novel weighted nonnegative least-squares framework, DeMixSC identifies and adjusts genes with high technological discrepancy and aligns the benchmark data with large patient cohorts of matched-tissue-type for large-scale deconvolution. Our results using a benchmark dataset of healthy retinas suggest much-improved deconvolution accuracy. Further analysis of a cohort of 453 patients with age-related macular degeneration supports the broad applicability of DeMixSC. Our findings reveal the impact of technological discrepancy on deconvolution performance and underscore the importance of a well-matched dataset to resolve this challenge. The developed DeMixSC framework is generally applicable for deconvolving large cohorts of disease tissues, and potentially cancer.