遇见数据集

"GermanFakeNC" - German Fake News Corpus including manually fact-checked false statements

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

"GermanFakeNC" is a German Fake News Corpus including 490 texts which were retrieved from German alternative online media sources. Every fake statement in the text was verifi ed claim-by-claim by authoritative sources (e.g. from local police authorities, scientific studies, the police press office, etc.). The time interval for most of the news is established from December 2015 to March 2018. Description of the .json file: Date: publication date of the article URL: URL of the website A maximum of three false statements are provided: False_Statement_[1-3]_Location: Location of the verified false statement - Title, Teaser or Text False_Statement_[1-3]_Index: The index numbers refer to the token (!) position / number. We tokenized the text with "spaCy" (the free open-source library for Python). Example: Title of Text: "The quick brown fox jumped over the lazy dog. The fox broke both his legs while jumping." False_Statement: The fox broke both his legs while jumping. "False_Statement_1_Location": "Title", "False_Statement_1_Index": "11-19" Ratio_of_Fake_Statements: Percentage of fake found in the article 1 = Text is based on true information. Up to 25% of the information in the text is false 2 = Up to 50% of the information in the text is false. The other statements in the article are factually accurate 3 = Up to 75% of the content non-factual and incorrect 4 = Pure fabrication with up to 100% false information in text 9 = Unclear, unverifiable Overall_Rating of the disinformation in text: range [0.1:1.0]. 0.1 no disinformation in text 0.2 0.3 0.4 0.5 neutral / ambivalent 0.6 0.7 0.8 0.9 1.0 strong disinformative text ====================================================================== The original sources retain the copyright of the data. You are allowed to use this dataset for research purposes only. For more question about the dataset, please contact: Inna Vogel, inna.vogel@sit.fraunhofer.de v1.01/09/2019

"GermanFakeNC"是一款德语假新闻语料库,收录了从德国非主流在线媒体来源抓取的490篇文本。文本中的每一条虚假陈述均由权威来源逐句核验,核验主体包括当地警方、学术研究机构、警方新闻办公室等。该数据集内多数新闻的发布时间跨度为2015年12月至2018年3月。 ### JSON文件字段说明 - `Date`:文章的发布日期 - `URL`:文章来源网页的链接 单篇文本最多标注三条虚假陈述: - `False_Statement_[1-3]_Location`:已核验虚假陈述的所在位置,可选值为标题(Title)、导语(Teaser)或正文(Text) - `False_Statement_[1-3]_Index`:索引值指代词元(Token)的位置与序号,本数据集使用Python开源免费库`spaCy`对文本进行词元化处理。 #### 示例: 文本标题:"The quick brown fox jumped over the lazy dog. The fox broke both his legs while jumping." 对应虚假陈述:"The fox broke both his legs while jumping." 标注格式如下: "False_Statement_1_Location": "Title", "False_Statement_1_Index": "11-19" - `Ratio_of_Fake_Statements`:文章虚假信息占比分级: 1级:文本主体基于真实信息,虚假信息占比不超过25% 2级:文本虚假信息占比不超过50%,其余陈述均符合事实 3级:文本中75%及以内的内容为非事实性错误信息 4级:完全虚构,文本虚假信息占比可达100% 9级:信息模糊,无法核验 - `Overall_Rating`:文本虚假信息整体评级,取值范围为[0.1, 1.0]: 0.1:文本无虚假误导信息 0.2、0.3、0.4:无额外说明 0.5:中立/态度模糊 0.6、0.7、0.8:无额外说明 0.9、1.0:文本存在极强的虚假误导性信息 ====================================================================== 本数据集的原始内容保留其原有版权。 本数据集仅可用于学术研究用途。 若需咨询数据集相关问题,请联系:英娜·沃格尔(Inna Vogel),邮箱:inna.vogel@sit.fraunhofer.de 版本v1.0,发布日期2019年9月1日。

创建时间:
2019-08-23
二维码
社区交流群
二维码
科研交流群
商业服务