<HashMap><database>biostudies-literature</database><scores/><additional><submitter>Alexander DL</submitter><funding>NIGMS NIH HHS</funding><pagination>1316-22</pagination><full_dataset_link>https://www.ebi.ac.uk/biostudies/studies/S-EPMC4530125</full_dataset_link><repository>biostudies-literature</repository><omics_type>Unknown</omics_type><volume>55(7)</volume><pubmed_abstract>The statistical metrics used to characterize the external predictivity of a model, i.e., how well it predicts the properties of an independent test set, have proliferated over the past decade. This paper clarifies some apparent confusion over the use of the coefficient of determination, R(2), as a measure of model fit and predictive power in QSAR and QSPR modeling. R(2) (or r(2)) has been used in various contexts in the literature in conjunction with training and test data for both ordinary linear regression and regression through the origin as well as with linear and nonlinear regression models. We analyze the widely adopted model fit criteria suggested by Golbraikh and Tropsha ( J. Mol. Graphics Modell. 2002 , 20 , 269 - 276 ) in a strict statistical manner. Shortcomings in these criteri</pubmed_abstract><journal>Journal of chemical information and modeling</journal><pubmed_title>Beware of R(2): Simple, Unambiguous Assessment of the Prediction Accuracy of QSAR and QSPR Models.</pubmed_title><pmcid>PMC4530125</pmcid><funding_grant_id>R01 GM096967</funding_grant_id><pubmed_authors>Winkler DA</pubmed_authors><pubmed_authors>Alexander DL</pubmed_authors><pubmed_authors>Tropsha A</pubmed_authors></additional><is_claimable>false</is_claimable><name>Beware of R(2): Simple, Unambiguous Assessment of the Prediction Accuracy of QSAR and QSPR Models.</name><description>The statistical metrics used to characterize the external predictivity of a model, i.e., how well it predicts the properties of an independent test set, have proliferated over the past decade. This paper clarifies some apparent confusion over the use of the coefficient of determination, R(2), as a measure of model fit and predictive power in QSAR and QSPR modeling. R(2) (or r(2)) has been used in various contexts in the literature in conjunction with training and test data for both ordinary linear regression and regression through the origin as well as with linear and nonlinear regression models. We analyze the widely adopted model fit criteria suggested by Golbraikh and Tropsha ( J. Mol. Graphics Modell. 2002 , 20 , 269 - 276 ) in a strict statistical manner. Shortcomings in these criteri</description><dates><release>2015-01-01T00:00:00Z</release><publication>2015 Jul</publication><modification>2025-04-05T00:03:58.183Z</modification><creation>2019-03-27T01:56:32Z</creation></dates><accession>S-EPMC4530125</accession><cross_references><pubmed>26099013</pubmed><doi>10.1021/acs.jcim.5b00206</doi></cross_references></HashMap>