<HashMap><database>biostudies-literature</database><scores/><additional><omics_type>Unknown</omics_type><submitter>Billato I</submitter><funding>NHGRI NIH HHS</funding><funding>NCI NIH HHS</funding><pubmed_abstract>The increasing size of single-cell RNA sequencing (scRNA-seq) datasets poses major computational challenges. This work benchmarks the scalability, efficiency, and accuracy of five widely used analysis frameworks (Seurat, OSCA, scrapper, Scanpy, and rapids_singlecell), focusing on the impact of algorithmic and infrastructural choices on performance. We performed a systematic comparison of these workflows using representative datasets, including a 1.3 million mouse brain cell dataset for scalability and three smaller datasets (BE1, scMixology, and cord blood CITE-seq) with ground truth labels to assess clustering accuracy. Principal Component Analysis (PCA) was used as a paradigmatic step to evaluate the computational performance of six SVD algorithms (exact, ARPACK, IRLBA, randomized, Jacob</pubmed_abstract><journal>bioRxiv : the preprint server for biology</journal><pagination>2025.10.28.681564</pagination><full_dataset_link>https://www.ebi.ac.uk/biostudies/studies/S-EPMC12636554</full_dataset_link><repository>biostudies-literature</repository><pubmed_title>Benchmarking large-scale single-cell RNA-seq analysis.</pubmed_title><pmcid>PMC12636554</pmcid><funding_grant_id>U24 CA289073</funding_grant_id><funding_grant_id>U24 HG004059</funding_grant_id><pubmed_authors>Carey V</pubmed_authors><pubmed_authors>Billato I</pubmed_authors><pubmed_authors>Waldron L</pubmed_authors><pubmed_authors>Romualdi C</pubmed_authors><pubmed_authors>Risso D</pubmed_authors><pubmed_authors>Pages H</pubmed_authors><pubmed_authors>Sales G</pubmed_authors></additional><is_claimable>false</is_claimable><name>Benchmarking large-scale single-cell RNA-seq analysis.</name><description>The increasing size of single-cell RNA sequencing (scRNA-seq) datasets poses major computational challenges. This work benchmarks the scalability, efficiency, and accuracy of five widely used analysis frameworks (Seurat, OSCA, scrapper, Scanpy, and rapids_singlecell), focusing on the impact of algorithmic and infrastructural choices on performance. We performed a systematic comparison of these workflows using representative datasets, including a 1.3 million mouse brain cell dataset for scalability and three smaller datasets (BE1, scMixology, and cord blood CITE-seq) with ground truth labels to assess clustering accuracy. Principal Component Analysis (PCA) was used as a paradigmatic step to evaluate the computational performance of six SVD algorithms (exact, ARPACK, IRLBA, randomized, Jacob</description><dates><release>2025-01-01T00:00:00Z</release><publication>2025 Oct</publication><modification>2026-06-11T03:13:03.751Z</modification><creation>2026-06-11T03:08:15.335Z</creation></dates><accession>S-EPMC12636554</accession><cross_references><pubmed>41279840</pubmed><doi>10.1101/2025.10.28.681564</doi></cross_references></HashMap>