{"database":"biostudies-literature","file_versions":[],"scores":null,"additional":{"omics_type":["Unknown"],"submitter":["Billato I"],"funding":["NHGRI NIH HHS","NCI NIH HHS"],"pubmed_abstract":["The increasing size of single-cell RNA sequencing (scRNA-seq) datasets poses major computational challenges. This work benchmarks the scalability, efficiency, and accuracy of five widely used analysis frameworks (Seurat, OSCA, scrapper, Scanpy, and rapids_singlecell), focusing on the impact of algorithmic and infrastructural choices on performance. We performed a systematic comparison of these workflows using representative datasets, including a 1.3 million mouse brain cell dataset for scalability and three smaller datasets (BE1, scMixology, and cord blood CITE-seq) with ground truth labels to assess clustering accuracy. Principal Component Analysis (PCA) was used as a paradigmatic step to evaluate the computational performance of six SVD algorithms (exact, ARPACK, IRLBA, randomized, Jacob"],"journal":["bioRxiv : the preprint server for biology"],"pagination":["2025.10.28.681564"],"full_dataset_link":["https://www.ebi.ac.uk/biostudies/studies/S-EPMC12636554"],"repository":["biostudies-literature"],"pubmed_title":["Benchmarking large-scale single-cell RNA-seq analysis."],"pmcid":["PMC12636554"],"funding_grant_id":["U24 CA289073","U24 HG004059"],"pubmed_authors":["Carey V","Billato I","Waldron L","Romualdi C","Risso D","Pages H","Sales G"],"additional_accession":[]},"is_claimable":false,"name":"Benchmarking large-scale single-cell RNA-seq analysis.","description":"The increasing size of single-cell RNA sequencing (scRNA-seq) datasets poses major computational challenges. This work benchmarks the scalability, efficiency, and accuracy of five widely used analysis frameworks (Seurat, OSCA, scrapper, Scanpy, and rapids_singlecell), focusing on the impact of algorithmic and infrastructural choices on performance. We performed a systematic comparison of these workflows using representative datasets, including a 1.3 million mouse brain cell dataset for scalability and three smaller datasets (BE1, scMixology, and cord blood CITE-seq) with ground truth labels to assess clustering accuracy. Principal Component Analysis (PCA) was used as a paradigmatic step to evaluate the computational performance of six SVD algorithms (exact, ARPACK, IRLBA, randomized, Jacob","dates":{"release":"2025-01-01T00:00:00Z","publication":"2025 Oct","modification":"2026-06-11T03:13:03.751Z","creation":"2026-06-11T03:08:15.335Z"},"accession":"S-EPMC12636554","cross_references":{"pubmed":["41279840"],"doi":["10.1101/2025.10.28.681564"]}}