Skip to main content
Have a personal or library account? Click to login
Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers Cover

Noor-Sharaye v.1. A Benchmark Dataset of Complex Words for Evaluating Arabic Analyzers

Open Access
|Aug 2026

Abstract

The Noor-Sharaye dataset is a morphologically annotated Classical Arabic corpus containing approximately 205,000 word instances extracted from 17 different Classical Arabic books, including Quranic, Fiqh, Hadith, and historical texts. Each token is enriched with detailed linguistic annotations such as stem, lemma, root, part-of-speech tags, segmentation, grammatical case, gender, number, and affix-level features, designed to capture the complex morphological structure of Classical Arabic. The data are encoded in UTF-8 XLS, XML, and JSON formats for broad compatibility. This resource supports stemming, root extraction, morphological analysis, and benchmarking of AI-based models in Arabic Natural Language Processing.

DOI: https://doi.org/10.5334/johd.572 | Journal eISSN: 2059-481X
Language: English
Page range: 110 - 110
Submitted on: Apr 19, 2026
Accepted on: Jul 24, 2026
Published on: Aug 13, 2026
Published by: Ubiquity Press
In partnership with: Paradigm Publishing Services

© 2026 Azal Alaswaad, Behrouz Minaei-Bidgoli, published by Ubiquity Press
This work is licensed under the Creative Commons Attribution 4.0 License.