遇见数据集

FakeSumm

收藏
arXiv2024-01-09 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

FakeSumm数据集由印度理工学院海得拉巴分校、帕蒙尼克、格拉斯哥大学和LDRP技术与研究学院合作创建,旨在通过对抗性提示技术生成用于检测错误信息的数据集。该数据集包含约5000条正确新闻摘要和1000条包含错误信息的新闻摘要,覆盖多种错误类型,如虚构、错误归属、不准确的量化数据和误导性陈述。创建过程中,研究者利用大型语言模型(如GPT-3.5和GPT-4)生成摘要,并通过精心设计的提示控制错误信息的类型和程度。FakeSumm数据集的应用领域主要集中在提高自动错误信息检测系统的性能,以应对日益增长的社交媒体错误信息问题。

The FakeSumm dataset was co-created by the Indian Institute of Technology Hyderabad, Paramonik, the University of Glasgow, and the LDRP Institute of Technology and Research, with the goal of generating datasets for misinformation detection via adversarial prompting techniques. It contains approximately 5,000 authentic news summaries and 1,000 misinformative news summaries, covering multiple types of misinformation such as fabrication, misattribution, inaccurate quantitative data, and misleading statements. During the development process, researchers used large language models (LLMs, e.g., GPT-3.5 and GPT-4) to generate the summaries, and controlled the type and severity of the embedded misinformation through meticulously designed prompts. The primary application of the FakeSumm dataset is to enhance the performance of automated misinformation detection systems to tackle the escalating problem of misinformation on social media.

创建时间:
2024-01-09
二维码
社区交流群
二维码
科研交流群
商业服务