coref-data/mmc_raw
收藏MMC (Multilingual Multiparty Coreference)
数据集配置
-
mmc_en:
- 训练集:
mmc_en/train-* - 开发集:
mmc_en/dev-* - 测试集:
mmc_en/test-*
- 训练集:
-
mmc_fa:
- 训练集:
mmc_fa/train-* - 开发集:
mmc_fa/dev-* - 测试集:
mmc_fa/test-*
- 训练集:
-
mmc_fa_corrected:
- 训练集:
mmc_fa_corrected/train-* - 开发集:
mmc_fa_corrected/dev-* - 测试集:
mmc_fa_corrected/test-*
- 训练集:
-
mmc_zh_corrected:
- 训练集:
mmc_zh_corrected/train-* - 开发集:
mmc_zh_corrected/dev-* - 测试集:
mmc_zh_corrected/test-*
- 训练集:
-
mmc_zh_uncorrected:
- 训练集:
mmc_zh_uncorrected/train-* - 开发集:
mmc_zh_uncorrected/dev-* - 测试集:
mmc_zh_uncorrected/test-*
- 训练集:
数据来源
- 数据集用于论文 "Multilingual Coreference Resolution in Multiparty Dialogue",发表于 TACL 2023。
引用
@article{zheng-etal-2023-multilingual, title = "Multilingual Coreference Resolution in Multiparty Dialogue", author = "Zheng, Boyuan and Xia, Patrick and Yarmohammadi, Mahsa and Van Durme, Benjamin", journal = "Transactions of the Association for Computational Linguistics", volume = "11", year = "2023", address = "Cambridge, MA", publisher = "MIT Press", url = "https://aclanthology.org/2023.tacl-1.52", doi = "10.1162/tacl_a_00581", pages = "922--940", abstract = "Existing multiparty dialogue datasets for entity coreference resolution are nascent, and many challenges are still unaddressed. We create a large-scale dataset, Multilingual Multiparty Coref (MMC), for this task based on TV transcripts. Due to the availability of gold-quality subtitles in multiple languages, we propose reusing the annotations to create silver coreference resolution data in other languages (Chinese and Farsi) via annotation projection. On the gold (English) data, off-the-shelf models perform relatively poorly on MMC, suggesting that MMC has broader coverage of multiparty coreference than prior datasets. On the silver data, we find success both using it for data augmentation and training from scratch, which effectively simulates the zero-shot cross-lingual setting.", }




