<HashMap><database>biostudies-literature</database><scores/><additional><omics_type>Unknown</omics_type><submitter>Tariq U</submitter><funding>NIGMS NIH HHS</funding><pubmed_abstract>Database search algorithms reduce the number of potential candidate peptides against which scoring needs to be performed using a single (i.e. mass) property for filtering. While useful, filtering based on one property may lead to exclusion of non-abundant spectra and uncharacterized peptides - potentially exacerbating the &lt;i>streetlight&lt;/i> effect. Here we present &lt;i>ProteoRift&lt;/i>, a novel attention and multitask deep-network, which can &lt;i>predict&lt;/i> multiple peptide properties (length, missed cleavages, and modification status) directly from spectra. We demonstrate that &lt;i>ProteoRift&lt;/i> can predict these properties with up to 97% accuracy resulting in search-space reduction by more than 90%. As a result, our end-to-end pipeline is shown to exhibit 8x to 12x speedups with peptide deduction accuracy comparable to algorithmic techniques. We also formulate two uncertainty estimation metrics, which can distinguish between in-distribution and out-of-distribution data (ROC-AUC 0.99) and predict high-scoring mass spectra against correct peptide (ROC-AUC 0.94). These models and metrics are integrated in an end-to-end ML pipeline available at https://github.com/pcdslab/ProteoRift.</pubmed_abstract><journal>bioRxiv : the preprint server for biology</journal><pagination>2024.08.21.609035</pagination><full_dataset_link>https://www.ebi.ac.uk/biostudies/studies/S-EPMC11370541</full_dataset_link><repository>biostudies-literature</repository><pubmed_title>Predicting peptide properties from mass spectrometry data using deep attention-based multitask network and uncertainty quantification.</pubmed_title><pmcid>PMC11370541</pmcid><funding_grant_id>R35 GM153434</funding_grant_id><pubmed_authors>Saeed F</pubmed_authors><pubmed_authors>Tariq U</pubmed_authors></additional><is_claimable>false</is_claimable><name>Predicting peptide properties from mass spectrometry data using deep attention-based multitask network and uncertainty quantification.</name><description>Database search algorithms reduce the number of potential candidate peptides against which scoring needs to be performed using a single (i.e. mass) property for filtering. While useful, filtering based on one property may lead to exclusion of non-abundant spectra and uncharacterized peptides - potentially exacerbating the &lt;i>streetlight&lt;/i> effect. Here we present &lt;i>ProteoRift&lt;/i>, a novel attention and multitask deep-network, which can &lt;i>predict&lt;/i> multiple peptide properties (length, missed cleavages, and modification status) directly from spectra. We demonstrate that &lt;i>ProteoRift&lt;/i> can predict these properties with up to 97% accuracy resulting in search-space reduction by more than 90%. As a result, our end-to-end pipeline is shown to exhibit 8x to 12x speedups with peptide deduction accuracy comparable to algorithmic techniques. We also formulate two uncertainty estimation metrics, which can distinguish between in-distribution and out-of-distribution data (ROC-AUC 0.99) and predict high-scoring mass spectra against correct peptide (ROC-AUC 0.94). These models and metrics are integrated in an end-to-end ML pipeline available at https://github.com/pcdslab/ProteoRift.</description><dates><release>2024-01-01T00:00:00Z</release><publication>2024 Aug</publication><modification>2026-07-02T03:20:19.768Z</modification><creation>2025-04-06T10:14:19.231Z</creation></dates><accession>S-EPMC11370541</accession><cross_references><pubmed>39229185</pubmed><doi>10.1101/2024.08.21.609035</doi></cross_references></HashMap>