Webis Wikipedia-IPC
收藏数据链接:
官方服务:
资源简介:
<strong>Webis Wikipedia-IPC</strong> When an image is reused on the Web, an original caption is often assigned. We hypothesize that different captions for the same image naturally form a set of mutual paraphrases. To demonstrate the suitability of this idea, we analyzed captions in the English Wikipedia, where editors frequently relabel the same image for different articles. As a result, the Wikipedia-IPC (<strong>I</strong>mage caption <strong>P</strong>araphrase <strong>C</strong>orpus) dataset was created which include caption pairs of the same image which represent paraphrases. It contains 30,237 gold, 229,877 silver, and 656,560 bronze quality paraphrase pairs.
创建时间:
2023-02-08



