AraLexSimp: Arabic Contextual Lexical Simplification Benchmark Dataset
收藏资源简介:
AraLexSimp is an Arabic benchmark dataset for contextual lexical simplification. It extends the AraLexSubD lexical substitution dataset by introducing human annotations for semantic preservation, lexical simplicity, and contextual simplicity. The benchmark contains Arabic sentences from three domains (General, Medical, and Financial), target complex words, candidate substitutions, substituted sentences, human evaluation scores, and gold-standard rankings for benchmarking contextual lexical simplification systems. Each candidate substitute was independently evaluated by three annotators using a five-point Likert scale. Semantic preservation was first used to filter unsuitable candidates, while contextual simplicity served as the primary ranking criterion. Lexical simplicity and semantic preservation were subsequently used to resolve ties. The dataset is intended to support the development, evaluation, and comparison of Arabic contextual lexical simplification methods and to facilitate reproducible research in Arabic Natural Language Processing (NLP). Repository contents:• AraLexSimp_Dataset.xlsx• README.md• DATA_DICTIONARY.md• LICENSE• CITATION.cff



