Dataset Information

A bilingual benchmark for evaluating large language models.

ABSTRACT: This work introduces a new benchmark for the bilingual evaluation of large language models (LLMs) in English and Arabic. While LLMs have transformed various fields, their evaluation in Arabic remains limited. This work addresses this gap by proposing a novel evaluation method for LLMs in both Arabic and English, allowing for a direct comparison between the performance of the two languages. We build a new evaluation dataset based on the General Aptitude Test (GAT), a standardized test widely used for university admissions in the Arab world, that we utilize to measure the linguistic capabilities of LLMs. We conduct several experiments to examine the linguistic capabilities of ChatGPT and quantify how much better it is at English than Arabic. We also examine the effect of changing task descriptions from Arabic to English and vice-versa. In addition to that, we find that fastText can surpass ChatGPT in finding Arabic word analogies. We conclude by showing that GPT-4 Arabic linguistic capabilities are much better than ChatGPT's Arabic capabilities and are close to ChatGPT's English capabilities.

SUBMITTER: Alkaoud M

PROVIDER: S-EPMC10909174 | biostudies-literature | 2024

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

A bilingual benchmark for evaluating large language models.

Alkaoud Mohamed M

PeerJ. Computer science 20240229

This work introduces a new benchmark for the bilingual evaluation of large language models (LLMs) in English and Arabic. While LLMs have transformed various fields, their evaluation in Arabic remains limited. This work addresses this gap by proposing a novel evaluation method for LLMs in both Arabic and English, allowing for a direct comparison between the performance of the two languages. We build a new evaluation dataset based on the General Aptitude Test (GAT), a standardized test widely used ...[more]

PMID: 38435597

Dataset Information

A bilingual benchmark for evaluating large language models.

Publications

A bilingual benchmark for evaluating large language models.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets

Similar Datasets

A benchmark dataset for evaluating gender sensitivity in Korean political discourse with large language models.
| S-EPMC12808084 | biostudies-literature

CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research: A novel question-and-answer benchmark designed to assess Large Language Models' comprehension of biomedical research, piloted on Neurodegenerative Diseases.
| S-EPMC11760394 | biostudies-literature

Benchmark evaluation of DeepSeek large language models in clinical decision-making.
| S-EPMC12353792 | biostudies-literature

Evaluating multiple large language models on orbital diseases.
| S-EPMC12277337 | biostudies-literature

Evaluating Large Language Models on Medical Evidence Summarization.
| S-EPMC10168498 | biostudies-literature

The Two Word Test as a semantic benchmark for large language models.
| S-EPMC11405709 | biostudies-literature

Evaluating large language models on multimodal chemistry olympiad exams.
| S-EPMC12717038 | biostudies-literature

Evaluating large language models in theory of mind tasks.
| S-EPMC11551352 | biostudies-literature

MedCalc-Bench: Evaluating Large Language Models for Medical Calculations.
| S-EPMC12478433 | biostudies-literature

Evaluating anti-LGBTQIA+ medical bias in large language models.
| S-EPMC12416741 | biostudies-literature