遇见数据集

AraLexSimp: Arabic Contextual Lexical Simplification Benchmark Dataset

收藏
Zenodo2026-07-11 更新2026-08-01 收录
官方服务:

资源简介:

AraLexSimp is an Arabic benchmark dataset for contextual lexical simplification. It extends the AraLexSubD lexical substitution dataset by introducing human annotations for semantic preservation, lexical simplicity, and contextual simplicity. The benchmark contains Arabic sentences from three domains (General, Medical, and Financial), target complex words, candidate substitutions, substituted sentences, human evaluation scores, and gold-standard rankings for benchmarking contextual lexical simplification systems. Each candidate substitute was independently evaluated by three annotators using a five-point Likert scale. Semantic preservation was first used to filter unsuitable candidates, while contextual simplicity served as the primary ranking criterion. Lexical simplicity and semantic preservation were subsequently used to resolve ties. The dataset is intended to support the development, evaluation, and comparison of Arabic contextual lexical simplification methods and to facilitate reproducible research in Arabic Natural Language Processing (NLP). Repository contents:• AraLexSimp_Dataset.xlsx• README.md• DATA_DICTIONARY.md• LICENSE• CITATION.cff

提供机构:
Zenodo
创建时间:
2026-07-11
二维码
社区交流群
二维码
科研交流群
商业服务