遇见数据集

ThaiCoref

收藏
arXiv2024-06-10 更新2024-06-12 收录
官方服务:

资源简介:

ThaiCoref是由朱拉隆功大学语言学系创建的泰语指代消解数据集,包含777,271个tokens、44,082个提及和10,429个实体,涵盖大学论文、报纸、演讲和维基百科四种文本类型。该数据集基于OntoNotes基准,针对泰语特定现象进行了调整。创建过程中,通过严格的标注指南和多轮校验确保数据质量,旨在解决泰语指代消解的挑战,支持泰语自然语言处理的研究和应用。

ThaiCoref is a Thai coreference resolution dataset created by the Department of Linguistics, Chulalongkorn University. It contains 777,271 tokens, 44,082 entity mentions, and 10,429 entities, covering four text types: university theses, newspapers, speeches, and Wikipedia. Built upon the OntoNotes benchmark, this dataset has been tailored to accommodate Thai-specific linguistic phenomena. During its development, strict annotation guidelines and multi-round validation were employed to ensure data quality. It aims to address the challenges of Thai coreference resolution and support research and applications in Thai natural language processing.

创建时间:
2024-06-10
搜集汇总
背景与挑战
背景概述
ThaiCoref是一个由朱拉隆功大学语言学系创建的泰语指代消解数据集,包含超过77万个tokens、4.4万个提及和1万个实体,覆盖大学论文、报纸、演讲和维基百科四种文本类型。它基于OntoNotes基准并针对泰语特点进行调整,通过严格标注和校验确保高质量,旨在支持泰语自然语言处理的研究和应用。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务