遇见数据集

Webis Wikipedia-IPC

收藏
Zenodo2023-04-28 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

<strong>Webis Wikipedia-IPC</strong> When an image is reused on the Web, an original caption is often assigned. We hypothesize that different captions for the same image naturally form a set of mutual paraphrases. To demonstrate the suitability of this idea, we analyzed captions in the English Wikipedia, where editors frequently relabel the same image for different articles. As a result, the Wikipedia-IPC (<strong>I</strong>mage caption <strong>P</strong>araphrase <strong>C</strong>orpus) dataset was created which include caption pairs of the same image which represent paraphrases. It contains 30,237 gold, 229,877 silver, and 656,560 bronze quality paraphrase pairs.

提供机构:
Zenodo
创建时间:
2023-04-28
二维码
社区交流群
二维码
科研交流群
商业服务