{"database":"biostudies-literature","file_versions":[],"scores":null,"additional":{"submitter":["Maurya NS"],"funding":["Science and Engineering Research Board"],"pagination":["14304"],"full_dataset_link":["https://www.ebi.ac.uk/biostudies/studies/S-EPMC8275802"],"repository":["biostudies-literature"],"omics_type":["Unknown"],"volume":["11(1)"],"pubmed_abstract":["Colorectal cancer (CRC) is a common cause of cancer-related deaths worldwide. The CRC mRNA gene expression dataset containing 644 CRC tumor and 51 normal samples from the cancer genome atlas (TCGA) was pre-processed to identify the significant differentially expressed genes (DEGs). Feature selection techniques Least absolute shrinkage and selection operator (LASSO) and Relief were used along with class balancing for obtaining features (genes) of high importance. The classification of the CRC dataset was done by ML algorithms namely, random forest (RF), K-nearest neighbour (KNN), and artificial neural networks (ANN). The significant DEGs were 2933, having 1832 upregulated and 1101 downregulated genes. The CRC gene expression dataset had 23,186 features. LASSO had performed better than Relie"],"journal":["Scientific reports"],"pubmed_title":["Transcriptome profiling by combined machine learning and statistical R analysis identifies TMEM236 as a potential novel diagnostic biomarker for colorectal cancer."],"pmcid":["PMC8275802"],"funding_grant_id":["SB/YS/LS-107/2014"],"pubmed_authors":["Maurya NS","Kushwaha S","Chawade A","Mani A"],"additional_accession":[]},"is_claimable":false,"name":"Transcriptome profiling by combined machine learning and statistical R analysis identifies TMEM236 as a potential novel diagnostic biomarker for colorectal cancer.","description":"Colorectal cancer (CRC) is a common cause of cancer-related deaths worldwide. The CRC mRNA gene expression dataset containing 644 CRC tumor and 51 normal samples from the cancer genome atlas (TCGA) was pre-processed to identify the significant differentially expressed genes (DEGs). Feature selection techniques Least absolute shrinkage and selection operator (LASSO) and Relief were used along with class balancing for obtaining features (genes) of high importance. The classification of the CRC dataset was done by ML algorithms namely, random forest (RF), K-nearest neighbour (KNN), and artificial neural networks (ANN). The significant DEGs were 2933, having 1832 upregulated and 1101 downregulated genes. The CRC gene expression dataset had 23,186 features. LASSO had performed better than Relie","dates":{"release":"2021-01-01T00:00:00Z","publication":"2021 Jul","modification":"2025-04-19T09:04:15.713Z","creation":"2022-02-10T20:01:01.173Z"},"accession":"S-EPMC8275802","cross_references":{"pubmed":["34253750"],"doi":["10.1038/s41598-021-92692-0"]}}