遇见数据集

COMPILING

收藏
arXiv2022-09-29 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

COMPILING数据集是由北京语言大学创建,旨在为汉语作为外语学习者提供复杂度可控的定义生成任务。该数据集包含127,757条记录,每条记录包括一个词、其定义、示例及两个复杂度测量。数据集通过结合《当代汉语学习词典》与《现代汉语词典》第七版构建,利用HSK词汇等级量化定义复杂度,适用于语言学习和辅助教学,尤其有助于低文化程度读者及语言障碍者。

The COMPILING dataset was developed by Beijing Language and Culture University to provide complexity-controllable definition generation tasks for learners of Chinese as a foreign language. It comprises 127,757 entries, each containing a target word, its associated definition, an illustrative example, and two complexity metrics. This dataset was compiled by integrating the *Contemporary Chinese Learning Dictionary* and the 7th edition of *Modern Chinese Dictionary*, and uses HSK vocabulary proficiency levels to quantify the complexity of definitions. It is applicable to language learning and supportive teaching, and is particularly beneficial for readers with limited educational backgrounds and individuals with language impairments.

提供机构:
北京语言大学
创建时间:
2022-09-29
搜集汇总
数据集介绍
COMPILING 数据集图片
背景与挑战
背景概述
COMPILING数据集是由北京语言大学创建的汉语定义生成数据集,包含127,757条记录,每条记录包括词、定义、示例和复杂度测量。它基于《当代汉语学习词典》和《现代汉语词典》第七版构建,通过HSK词汇等级量化定义复杂度,适用于汉语作为外语的学习和教学,特别对低文化程度读者和语言障碍者有辅助作用。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务