GOLEMcoref multilingual annotated corpus
收藏DataCite Commons2026-05-06 更新2026-05-07 收录
下载链接:
https://zenodo.org/doi/10.5281/zenodo.20058309
下载链接
链接失效反馈官方服务:
资源简介:
A multilingual coreference dataset of 827k tokens of fiction in 7 languages: Bahasa Indonesia, Chinese, Dutch, English, Italian, Korean, and Spanish. The dataset includes full stories of diverse lengths, ranging from 500 to 17k words. The dataset is available under a creative commons license in CoNLL-2012 and CorefUD format.
提供机构:
Zenodo
创建时间:
2026-05-06



