Datasets for the paper: Lost in Translation: Using Global Fact-Checks to Measure Multilingual Misinformation Prevalence, Spread, and Evolution
收藏资源简介:
FullData.csv.gz: Contains links to all claims in the data-set. publishing_date: Date on which the fact-check was published. claim_date: Date that claim was made. verdict: Rating given by the fact-checking organisation. language: Language of the claim. cluster_{threshold}: ID of the cluster that claim belongs to at all given clusters. Entry "0" means that claim is singleton and not clustered with any other claims. Embeddings.npy: Contains a dictionary linking each claim to it's embedding calculated with LaBSE.
FullData.csv.gz:包含数据集中所有断言(claim)的链接。 publishing_date:该事实核查报告的发布日期。 claim_date:该断言的提出日期。 verdict:事实核查机构给出的评级结果。 language:该断言所使用的语言。 cluster_{threshold}:该断言在所有给定聚类阈值下所属聚类的ID。若条目为"0",则表示该断言为孤立样本,未与任何其他断言聚类。 Embeddings.npy:包含一个字典,将每个断言与其通过LaBSE计算得到的嵌入向量相关联。



