MegaFake
收藏资源简介:
MegaFake数据集由香港理工大学开发,是一个包含由大型语言模型生成的新闻数据的全面数据集,涵盖四种虚假新闻和两种合法新闻。该数据集基于GossipCop数据集构建,包含46,096条虚假新闻和17,871条合法新闻,是首个公开可用的大型机器生成虚假新闻数据集。MegaFake数据集的创建过程结合了社会心理学理论,通过一个新颖的自动化流水线生成,无需手动标注。该数据集主要用于支持未来在大型语言模型时代对虚假新闻检测和治理的研究。
The MegaFake dataset, developed by The Hong Kong Polytechnic University, is a comprehensive collection of news data generated by large language models (LLMs), encompassing four categories of fake news and two categories of legitimate news. Built upon the GossipCop dataset, it contains 46,096 fake news samples and 17,871 legitimate news samples, making it the first publicly available large-scale machine-generated fake news dataset. The development of the MegaFake dataset integrates social psychology theories and is generated through a novel automated pipeline without the need for manual annotation. This dataset is primarily designed to support future research on fake news detection and governance in the era of large language models.

- 1MegaFake: A Theory-Driven Dataset of Fake News Generated by Large Language Models香港理工大学 · 2024年



