MultiCaption: Dataset for detecting disinformation using multilingual visual claims
收藏资源简介:
MultiCaption is a multilingual dataset designed to identify disinformation through contradictory visual claims. Pairs of claims referring to the same image or video were labeled through multiple strategies to determine whether they contradict each other. The resulting dataset comprises 11,088 visual claims in 64 languages, offering a unique resource for building and evaluating misinformation-detection systems in truly multimodal and multilingual environments. The dataset contains two splits, train and test, as follows: No of. Contradicting Pairs No of. Non-contradicting Pairs Total Pairs No of. Languages Train 4020 4767 8795 59 Test 2415 2505 4920 52 Content: cid_1, cid_2 - Factchecked claim IDs from the original dataset MultiClaim v2 claim_1, claim_2 - Claims in their original language claim_1_en, claim_2_en, Claims in English translation type_1, type_2 - Type of the claim indicating claim/title/synthetic_title. claim refers to the original claim, and title refers to the claim retrieved from its corresponding fact-checked article title. synthetic_title or synthetic_claim refers to the paraphrase generated using GPT5 from its title or claim, respectively. language_1, language_2 - Language of the claims label_name - Label indicating contradicting or non-contradicting label - 1/0 contradicting or non-contradicting label_strategy - Strategy used for annotation. This includes Manual/Self-Expansion/Claim-Pair-Link/LLM-Annotation/GP5-paraphrase Preprint: https://arxiv.org/abs/2601.11220 References If you use MultiCaption in any publication, project, tool, or in any other form, please cite the following paper: @misc{frade2026multicaption, title={MultiCaption: Detecting disinformation using multilingual visual claims}, author={Rafael Martins Frade and Rrubaa Panchendrarajan and Arkaitz Zubiaga}, year={2026}, eprint={2601.11220}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2601.11220}, }



