Table 1
Datasets stemming.
| DATASET | WORD COUNT | LANGUAGE | IMPLEMENTATION |
|---|---|---|---|
| Noor-Sharaye Dataset | 205,000 | Classical texts | |
| CBAS Dataset | 20,291 papers | MSA (Modern standard Arabic) | (El-Defrawy, El-Sonbaty & Belal, 2015) |
| ICA (International Corpus of Arabic) dataset | 100 million words | MSA | (Alansary, Nagi & Adly, 2007) |
Table 2
Overview of Arabic Stemming and NLP Tools.
| TOOL | TYPE | MAIN APPROACH | KEY FEATURES | IMPLEMENTATION |
|---|---|---|---|---|
| Khoja Stemmer | Root-based stemmer | Removes prefixes/suffixes + pattern matching | Uses lexicons, stopwords, root/stem dictionaries | (Khoja & Garside, 1999) |
| Light-10 Stemmer | Light stemmer | Affix stripping + normalization | Handles common prefixes/suffixes, removes stopwords | (Soudi, Neumann & van den Bosch, 2007) |
| Tashaphyne | Light stemmer | Rule + lookup-based stemming | Diacritics removal, affix segmentation | (Al-Khatib et al., 2023) |
| NLTK | NLP toolkit | Multi-functional NLP library | Tokenization, stemming, parsing, classification | (Elbes et al., 2019) |
| P-Stemmer | Light stemmer | Improved light stemming | Removes only prefixes for better accuracy | (Kanan et al., 2019) |
| CAMeL Tools | NLP framework | Full Arabic NLP pipeline | Tokenization, tagging, sentiment, morphology | (Obiad, 2020) |
| Al-Khalil Analyzer | Morphological analyzer | Rule-based morphological parsing | Root extraction, diacritization support | (Boudlal et al., 2010) |
Table 3
Comparison of Noor-Sharaye Dataset with Existing Arabic NLP Resources.
| RESOURCE | TYPE | LANGUAGE VARIETY | MAIN FEATURES | IMPLEMENTATION |
|---|---|---|---|---|
| Noor Sharaye Dataset | Morphologically Annotated Corpus | Classical (17 sources) | Full morphology + stemming + benchmarking | |
| Quranic Arabic Corpus | Annotated Corpus | Classical Arabic (Quranic Arabic) | Lemma, POS, Root | (Dukes & Habash, 2010) |
| MADAMIRA | Morphological Analyzer and Disambiguator | Modern Standard Arabic and Arabic Dialects | Full morphology | (Pasha et al., 2014) |
| Farasa | Arabic NLP Toolkit | Modern Standard Arabic | Segmentation, stemming, POS tagging, named entity recognition | (Ahmed Abdelali et al., 2016) |
| Buckwalter Analyzer | Morphological Analyzer | Modern Standard Arabic | Morphological analysis | (Buckwalter, 2004) |

Figure 1
Sample of the dataset in xls format.
Table 4
Main Fields of the Noor-sharaye Dataset.
| FIELD | DESCRIPTION |
|---|---|
| Entry | Surface word form |
| Stem | Stem extracted from the word |
| Lemma | dictionary form |
| Root | Arabic root of the word |
| POS | Part-of-speech tag |
| Affix | Prefixes and suffixes associated with the word |
| Gender | Grammatical gender |
| Number | Singular, dual, or plural |
| Case | Grammatical case information |
| Categ | Morphological category |
Table 5
Presents a sample morphological annotation from the dataset.
| SURFACE FORM | LEMMA | ROOT | POS | SEGMENTATION | MORPHOLOGICAL FEATURES |
|---|---|---|---|---|---|
| فَسَأَلْتُمُونِيهَا (You asked Me for it.) | سَأَلَ | س-أ-ل | VERB | فَ + سَأَلْ +تُمُو + نِي + هَا | Tense: Past, Person:2, Gender:Masc, Number: Plural, Voice: Active |
| عِبَادَتُكُمْ (Your worship) | عِبَادَة | ع-ب-د | NOUN | عِبَادَة + كُمْ | Case: Nominative, Number: Singular, Gender: Feminine |
| بِالْحَقِّ (With truth) | حَقّ | ح-ق-ق | NOUN | بِ + الْ + حَقِّ | Case: Genitive, Number: Singular, Gender: Masculine, Definite |

Figure 2
Sample records from the released Noor-Sharaye dataset.

Figure 3
Example Python code for loading the dataset.
