Unknown

Dataset Information

0

Dataset of Karakalpak language stop words.


ABSTRACT: The dataset presented in this paper aims to address the challenge of automatic extraction of stop words in Natural Language Processing (NLP) for the low-resource Karakalpak language spoken by approximately two million people in Uzbekistan. To accomplish this, we have created a corpus of 23 Karakalpak language school textbooks, which we have named the Karakalpak Language School Corpus (KAASC). Using the KAASC corpus, we have constructed lists of stop words using three methods based on Term Frequency-Inverse Document Frequency (TF-IDF): unigram, bigram, and collocation methods, respectively. The resulting lists of stop words, along with a list of URLs used to construct the corpus, make up the described dataset in this paper.

SUBMITTER: Madatov K 

PROVIDER: S-EPMC10126844 | biostudies-literature | 2023 Jun

REPOSITORIES: biostudies-literature

altmetric image

Publications

Dataset of Karakalpak language stop words.

Madatov Khabibulla K   Bekchanov Shukurla S   Vičič Jernej J  

Data in brief 20230405


The dataset presented in this paper aims to address the challenge of automatic extraction of stop words in Natural Language Processing (NLP) for the low-resource Karakalpak language spoken by approximately two million people in Uzbekistan. To accomplish this, we have created a corpus of 23 Karakalpak language school textbooks, which we have named the Karakalpak Language School Corpus (KAASC). Using the KAASC corpus, we have constructed lists of stop words using three methods based on Term Freque  ...[more]

Similar Datasets

| S-EPMC10439288 | biostudies-literature
| S-EPMC9679746 | biostudies-literature
| S-EPMC7689026 | biostudies-literature
| S-EPMC7378574 | biostudies-literature
| S-EPMC6746567 | biostudies-literature
| S-EPMC9679712 | biostudies-literature
| S-EPMC8530275 | biostudies-literature
| S-EPMC3666749 | biostudies-literature
| S-EPMC10964063 | biostudies-literature
| S-EPMC10790027 | biostudies-literature