遇见数据集

GOLEMcoref multilingual annotated corpus

收藏
Zenodo2026-05-17 更新2026-05-26 收录
官方服务:

资源简介:

A multilingual coreference dataset of 827k tokens of fiction in 7 languages: Bahasa Indonesia, Chinese, Dutch, English, Italian, Korean, and Spanish. The dataset includes full stories of diverse lengths, ranging from 500 to 17k words. The dataset is available under a creative commons license in CoNLL-2012 and CorefUD format.

提供机构:
Zenodo
创建时间:
2026-05-06
二维码
社区交流群
二维码
科研交流群
商业服务