SM-D, AIGTBench
收藏资源简介:
SM-D数据集由香港科技大学(广州)和CISPA亥姆霍兹信息安全中心的研究团队创建,旨在量化社交媒体平台上AI生成文本(AIGT)的普及情况。该数据集包含了来自Medium、Quora和Reddit三个平台的约240万条帖子,时间跨度为2022年1月至2024年10月。AIGTBench数据集则是一个用于训练和评估AIGT检测器的基准数据集,包含了由12个不同的大型语言模型生成的约2877万条AIGT样本和1355万条HWT样本。AIGTBench的创建过程结合了开源数据集和基于社交媒体文本生成的AIGT数据,旨在为AIGT检测器提供多样化的训练和评估环境。该数据集的应用领域主要集中在社交媒体内容的AI生成文本检测,旨在解决AIGT在社交媒体上的滥用问题,如虚假信息传播和舆论操纵。
The SM-D dataset was developed by research teams from The Hong Kong University of Science and Technology (Guangzhou) and CISPA Helmholtz Center for Information Security, aiming to quantify the prevalence of AI-generated text (AIGT) on social media platforms. This dataset contains approximately 2.4 million posts from three platforms: Medium, Quora, and Reddit, spanning from January 2022 to October 2024. The AIGTBench dataset is a benchmark dataset for training and evaluating AIGT detectors, containing around 28.77 million AIGT samples and 13.55 million human-written text (HWT) samples generated by 12 distinct large language models (LLMs). The development of AIGTBench combines open-source datasets and AIGT data generated from social media text, aiming to provide diverse training and evaluation environments for AIGT detectors. The main application scenarios of this dataset focus on AI-generated text detection for social media content, aiming to address the abuse of AIGT on social media, such as the spread of disinformation and public opinion manipulation.

- 1Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media香港科技大学(广州), CISPA亥姆霍兹信息安全中心 · 2024年



