H-Prop and H-Prop-News Propaganda Datasets in Hindi
收藏资源简介:
The H-Prop dataset contains 28,630 articles created by translating a portion of Proppy Corpus in Hindi. Each article is labeled as either “propagandistic” (positive class) or “non-propagandistic” (negative class). The labeling done indirectly in Proppy corpus using a technique known as distant supervision is retained. The H-Prop-News dataset contains 5,500 Hindi News articles collected from 30+ prominent Hindi News websites. Each article is labeled as either “propagandistic” (positive class) or “non-propagandistic” (negative class). The labeling was done by human annotators and the inter-annotator agreement using Cohen’s Kappa measure observed is 0.81. ## Data format We provide the H-Prop dataset in three tsv files, including training, testing and validation partitions. The H-Prop-News dataset is provided in csv files including training, testing and validation partitions. Each line represents one article in H-Prop dataset with the following information: 1. article_text: the text of the article translated from Proppy corpus.<br> 2. propaganda_label: label for articles retained from Proppy corpus. Each line represents one article in H-Prop-News dataset with the following information: 1. news_website: Name of the news source website<br> 2. article_url: the direct URL for the published article in its source website<br> 3. news_headline: news headline<br> 4. article_text: the text of the article retrieved via parsehub tool<br> 5. propaganda_label: label for articles ## About The H-Prop dataset was translated using IBM Watson Language Translator. ## Credit Please cite the dataset as:<br> [HProp-News] Deptii Chaudhari, Ambika Pawar, and Alberto Barrón-Cedeño. 2022. H-Prop and H-Prop-News: Computational Propaganda Datasets in Hindi. doi: 10.5281/zenodo.5828240 ## Authors Deptii Chaudhari;<br> Ambika Pawar;<br> Alberto Barrón-Cedeno
H-Prop数据集(H-Prop dataset)包含28630篇文章,均由印地语Proppy语料库(Proppy Corpus)的部分内容翻译而来。每篇文章均标注为“宣传类(propagandistic,正类)”或“非宣传类(non-propagandistic,负类)”,我们保留了Proppy语料库中采用远监督(distant supervision)技术间接完成的标注规则。 H-Prop-News数据集(H-Prop-News dataset)包含5500篇印地语新闻文章,采集自30余家知名印地语新闻网站。每篇文章同样标注为“宣传类”或“非宣传类”,标注工作由人类标注员完成,经科恩卡帕系数(Cohen’s Kappa)计算得到的标注者间一致性为0.81。 ### 数据格式 我们将H-Prop数据集以TSV文件(TSV files)的形式提供,包含训练集、测试集与验证集三个数据划分。H-Prop-News数据集则以CSV文件(CSV files)形式提供,同样包含上述三个数据划分。 H-Prop数据集的每一行对应一篇文章,包含以下信息: 1. article_text:从Proppy语料库翻译而来的文章文本 2. propaganda_label:源自Proppy语料库的文章标注标签 H-Prop-News数据集的每一行对应一篇文章,包含以下信息: 1. news_website:新闻来源网站名称 2. article_url:来源网站上已发布文章的直接URL 3. news_headline:新闻标题 4. article_text:通过ParseHub工具获取的文章文本 5. propaganda_label:文章标注标签 ### 数据集说明 H-Prop数据集通过IBM Watson语言翻译器(IBM Watson Language Translator)完成翻译。 ### 引用说明 请按如下方式引用该数据集: [HProp-News] Deptii Chaudhari, Ambika Pawar, and Alberto Barrón-Cedeño. 2022. H-Prop and H-Prop-News: Computational Propaganda Datasets in Hindi. doi: 10.5281/zenodo.5828240 ### 作者 Deptii Chaudhari; Ambika Pawar; Alberto Barrón-Cedeno



