On the Role of Images for Analyzing Claims in Social Media
收藏资源简介:
This is a multimodal dataset used in the paper "On the Role of Images for Analyzing Claims in Social Media", accepted at CLEOPATRA-2021 (2nd International Workshop on Cross-lingual Event-centric Open Analytics), co-located with The Web Conference 2021.<br> <br> The four datasets are curated for two different tasks that broadly come under fake news detection. Originally, the datasets were released as part of challenges or papers for text-based NLP tasks and are further extended here with corresponding images. 1. <em><strong>clef_en</strong></em> and <em><strong>clef_ar</strong></em> are English and Arabic Twitter datasets for claim check-worthiness detection released in CLEF CheckThat! 2020 <em>Barr</em>ó<em>n-Cedeno et al.</em><sup> [1]</sup>.<br> 2. <strong><em>lesa</em></strong> is an English Twitter dataset for claim detection released by <em>Gupta et al.</em><sup>[2]</sup><br> 3. <strong><em>mediaeval</em></strong><em> </em>is an English Twitter dataset for conspiracy detection released in MediaEval 2020 Workshop by <em>Pogorelov et al.</em><sup>[3]</sup> The dataset details like data curation and annotation process can be found in the cited papers.<br> <br> Datasets released here with corresponding images are relatively smaller than the original text-based tweets. The data statistics are as follows:<br> 1. <strong><em>clef_en</em></strong>: 281<br> 2. <em><strong>clef_ar</strong></em>: 2571<br> 3. <strong><em>lesa</em></strong>: 1395<br> 4. <em><strong>mediaeval</strong></em>: 1724<br> <br> Each folder has two sub-folders and a json file <em>data.json</em> that consists of crawled tweets. Two sub-folders are:<br> 1. <em>images</em>: This Contains crawled images with the same name as tweet-id in <em>data.json</em>.<br> 2. <em>splits</em>: This contains 5-fold splits used for training and evaluation in our paper. Each file in this folder is a csv with two columns <em><tweet-id, label>.</em> Code for the paper: https://github.com/cleopatra-itn/image_text_claim_detection If you find the dataset and the paper useful, please cite our paper and the corresponding dataset papers<sup>[1,2,3]</sup><br> <strong>Cheema, Gullal S., et al. "On the Role of Images for Analyzing Claims in Social Media" </strong><em>2<sup>nd</sup> International Workshop on Cross-lingual Event-centric Open Analytics (CLEOPATRA) co-located with The Web Conf 2021.</em> [1] Barrón-Cedeno, Alberto, et al. "Overview of CheckThat! 2020: Automatic identification and verification of claims in social media." <em>International Conference of the Cross-Language Evaluation Forum for European Languages</em>. Springer, Cham, 2020.<br> [2] Gupta, Shreya, et al. "LESA: Linguistic Encapsulation and Semantic Amalgamation Based Generalised Claim Detection from Online Content." <em>arXiv preprint arXiv:2101.11891</em> (2021).<br> [3] Pogorelov, Konstantin, et al. "FakeNews: Corona Virus and 5G Conspiracy Task at MediaEval 2020." <em>MediaEval 2020 Workshop</em>. 2020.



