Unknown

Dataset Information

0

The model student: GPT-4 performance on graduate biomedical science exams.


ABSTRACT: The GPT-4 large language model (LLM) and ChatGPT chatbot have emerged as accessible and capable tools for generating English-language text in a variety of formats. GPT-4 has previously performed well when applied to questions from multiple standardized examinations. However, further evaluation of trustworthiness and accuracy of GPT-4 responses across various knowledge domains is essential before its use as a reference resource. Here, we assess GPT-4 performance on nine graduate-level examinations in the biomedical sciences (seven blinded), finding that GPT-4 scores exceed the student average in seven of nine cases and exceed all student scores for four exams. GPT-4 performed very well on fill-in-the-blank, short-answer, and essay questions, and correctly answered several questions on figures sourced from published manuscripts. Conversely, GPT-4 performed poorly on questions with figures containing simulated data and those requiring a hand-drawn answer. Two GPT-4 answer-sets were flagged as plagiarism based on answer similarity and some model responses included detailed hallucinations. In addition to assessing GPT-4 performance, we discuss patterns and limitations in GPT-4 capabilities with the goal of informing design of future academic examinations in the chatbot era.

SUBMITTER: Stribling D 

PROVIDER: S-EPMC10920673 | biostudies-literature | 2024 Mar

REPOSITORIES: biostudies-literature

altmetric image

Publications

The model student: GPT-4 performance on graduate biomedical science exams.

Stribling Daniel D   Xia Yuxing Y   Amer Maha K MK   Graim Kiley S KS   Mulligan Connie J CJ   Renne Rolf R  

Scientific reports 20240307 1


The GPT-4 large language model (LLM) and ChatGPT chatbot have emerged as accessible and capable tools for generating English-language text in a variety of formats. GPT-4 has previously performed well when applied to questions from multiple standardized examinations. However, further evaluation of trustworthiness and accuracy of GPT-4 responses across various knowledge domains is essential before its use as a reference resource. Here, we assess GPT-4 performance on nine graduate-level examination  ...[more]

Similar Datasets

| S-EPMC7145268 | biostudies-literature
| S-EPMC4353078 | biostudies-literature
| S-EPMC8445460 | biostudies-literature
| S-EPMC8670685 | biostudies-literature
| S-EPMC6755307 | biostudies-literature
| S-EPMC10099880 | biostudies-literature
| S-EPMC6755206 | biostudies-literature