Skip to main content
Have a personal or library account? Click to login
Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers Cover

Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers

Open Access
|Aug 2026

Figures & Tables

Table 1

Datasets stemming.

DATASETWORD COUNTLANGUAGEIMPLEMENTATION
Noor-Sharaye Dataset205,000Classical texts
CBAS Dataset20,291 papersMSA (Modern standard Arabic)(El-Defrawy, El-Sonbaty & Belal, 2015)
ICA (International Corpus of Arabic) dataset100 million wordsMSA(Alansary, Nagi & Adly, 2007)
Table 2

Overview of Arabic Stemming and NLP Tools.

TOOLTYPEMAIN APPROACHKEY FEATURESIMPLEMENTATION
Khoja StemmerRoot-based stemmerRemoves prefixes/suffixes + pattern matchingUses lexicons, stopwords, root/stem dictionaries(Khoja & Garside, 1999)
Light-10 StemmerLight stemmerAffix stripping + normalizationHandles common prefixes/suffixes, removes stopwords(Soudi, Neumann & van den Bosch, 2007)
TashaphyneLight stemmerRule + lookup-based stemmingDiacritics removal, affix segmentation(Al-Khatib et al., 2023)
NLTKNLP toolkitMulti-functional NLP libraryTokenization, stemming, parsing, classification(Elbes et al., 2019)
P-StemmerLight stemmerImproved light stemmingRemoves only prefixes for better accuracy(Kanan et al., 2019)
CAMeL ToolsNLP frameworkFull Arabic NLP pipelineTokenization, tagging, sentiment, morphology(Obiad, 2020)
Al-Khalil AnalyzerMorphological analyzerRule-based morphological parsingRoot extraction, diacritization support(Boudlal et al., 2010)
Table 3

Comparison of Noor-Sharaye Dataset with Existing Arabic NLP Resources.

RESOURCETYPELANGUAGE VARIETYMAIN FEATURESIMPLEMENTATION
Noor Sharaye DatasetMorphologically Annotated CorpusClassical (17 sources)Full morphology + stemming + benchmarking
Quranic Arabic CorpusAnnotated CorpusClassical Arabic (Quranic Arabic)Lemma, POS, Root(Dukes & Habash, 2010)
MADAMIRAMorphological Analyzer and DisambiguatorModern Standard Arabic and Arabic DialectsFull morphology(Pasha et al., 2014)
FarasaArabic NLP ToolkitModern Standard ArabicSegmentation, stemming, POS tagging, named entity recognition(Ahmed Abdelali et al., 2016)
Buckwalter AnalyzerMorphological AnalyzerModern Standard ArabicMorphological analysis(Buckwalter, 2004)
Figure 1

Sample of the dataset in xls format.

Table 4

Main Fields of the Noor-sharaye Dataset.

FIELDDESCRIPTION
EntrySurface word form
StemStem extracted from the word
Lemmadictionary form
RootArabic root of the word
POSPart-of-speech tag
AffixPrefixes and suffixes associated with the word
GenderGrammatical gender
NumberSingular, dual, or plural
CaseGrammatical case information
CategMorphological category
Table 5

Presents a sample morphological annotation from the dataset.

SURFACE FORMLEMMAROOTPOSSEGMENTATIONMORPHOLOGICAL FEATURES
فَسَأَلْتُمُونِيهَا (You asked Me for it.)سَأَلَس-أ-لVERBفَ + سَأَلْ +تُمُو + نِي + هَاTense: Past, Person:2, Gender:Masc, Number: Plural, Voice: Active
عِبَادَتُكُمْ (Your worship)عِبَادَةع-ب-دNOUNعِبَادَة + كُمْCase: Nominative, Number: Singular, Gender: Feminine
بِالْحَقِّ (With truth)حَقّح-ق-قNOUNبِ + الْ + حَقِّCase: Genitive, Number: Singular, Gender: Masculine, Definite
Figure 2

Sample records from the released Noor-Sharaye dataset.

Figure 3

Example Python code for loading the dataset.

DOI: https://doi.org/10.5334/johd.572 | Journal eISSN: 2059-481X
Language: English
Page range: 110 - 110
Submitted on: Apr 19, 2026
Accepted on: Jul 24, 2026
Published on: Aug 13, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Azal Alaswaad, Behrouz Minaei-Bidgoli, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.