JaLeCoN
收藏资源简介:
JaLeCoN是由奈良先端科学技术大学院大学创建的日语词汇复杂度数据集,专为非母语读者设计。该数据集包含18220条数据,涵盖单个词汇及多词表达,并提供中文/韩文注释者和其他注释者的独立复杂度评分,以满足不同母语背景读者的需求。数据集的创建过程包括从新闻和政府文件中提取文本,并进行细致的词汇分割和复杂度评分。JaLeCoN主要用于日语词汇复杂度预测研究,旨在帮助开发辅助阅读工具,提高非母语读者的阅读理解能力。
JaLeCoN is a Japanese lexical complexity dataset developed by Nara Institute of Science and Technology, designed exclusively for non-native readers of Japanese. This dataset contains 18,220 entries covering both single lexical items and multi-word expressions, and provides independent complexity ratings from annotators with Chinese or Korean native language backgrounds as well as other annotators to meet the needs of readers with diverse native language backgrounds. The dataset was constructed by extracting texts from news articles and government documents, followed by meticulous lexical segmentation and complexity rating work. JaLeCoN is primarily used for research on Japanese lexical complexity prediction, aiming to assist in the development of reading assistance tools to improve the reading comprehension abilities of non-native Japanese readers.




