
Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers
Abstract
The Noor-Sharaye dataset is a morphologically annotated Classical Arabic corpus containing approximately 205,000 word instances extracted from 17 different Classical Arabic books, including Quranic, Fiqh, Hadith, and historical texts. Each token is enriched with detailed linguistic annotations such as stem, lemma, root, part-of-speech tags, segmentation, grammatical case, gender, number, and affix-level features, designed to capture the complex morphological structure of Classical Arabic. The data are encoded in UTF-8 XLS, XML, and JSON formats for broad compatibility. This resource supports stemming, root extraction, morphological analysis, and benchmarking of AI-based models in Arabic Natural Language Processing.
© 2026 Azal Alaswaad, Behrouz Minaei-Bidgoli, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.