遇见数据集

Dataset on Tracing Artificial Campaigns in Social Media Streams

收藏
Zenodo2026-05-25 更新2026-05-26 收录
官方服务:

资源简介:

The data set contains human and GPT generated social media posts (tweets) on multiple topics and configurations. It has been used to fintune / train RoBERTa and SNN-based detectors for artificial textual content. The data is currently part of a publication process, more information will be provided with the published work soon. In a nutshell, the provided data sets contains: Training data: 17,000 (A)/16,000 (H) Artificial posts on multiple topics were generated using the OpenAI GPT3.5 API (fewshot) 20,000 (A)/20,000 (H) Artificial posts on multiple topics were generated using the OpenAI GPT3.5 API (zeroshot) Posts containing fewer than 30 tokens (based on the GPT-3.5 tokenizer) were removed during analysis, resulting in 16,765 human-written and 15,692 artificial posts (fewshot) as well as 16,896 human-written and 15,082 artificial posts (zeroshot), respectively. Test data: 5,000 (A)/5,000 (H) Test dataset to show transferability; artificial & human posts (topic: Trump-related). 10,000 (A)/10,000 (H) Test dataset for transferability; also for threshold setting; artificial & human posts (topic: climate-related). Artificial Campaigns (Campaign stereotypes that consist of LLM-gemerated content): 125 (A)/9362 (H) Dataset comprising campaign pattern with a burst of 125 artificial posts (many accounts). 115 (A)/12966 (H) Dataset with overlapping campaign patterns by few accounts over time. 100 (A)/8847 (H) Dataset’s campaign patterns result from initially strong but fading account activity over time. For training RoBERTa use the following code provided by Yang, Lingyi and Jiang, Feng and Li, Haizhou (2023): https://github.com/FreedomIntelligence/ChatGPT-Detection-PR-HPPT For stream data analysis, use the TextClust implementation provided by Assenmacher et al. (2022): https://github.com/Dennis1989/textClust-experiments For FastDetectGPT we used the implementation by Bao et al (2023): https://github.com/baoguangsheng/fast-detect-gpt References: Assenmacher, Dennis and Trautmann, Heike (2022). Textual one-pass stream clustering with automated distance threshold adaption. In: Asian conference on intelligent information and database systems. Springer; 2022. p. 3–16. Bao, Guangsheng and Zhao, Yanbin and Teng, Zhiyang and Yang, Linyi and Zhang, Yue (2023). Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature. The Twelfth International Conference on Learning Representations. Yang, Lingyi and Jiang, Feng and Li, Haizhou (2023). Is chatgpt involved in texts? Measure the polish ratio to detect chatgpt-generated text. APSIPA Transactions on Signal and Information Processing. 13(2).

提供机构:
Zenodo
创建时间:
2026-05-25
二维码
社区交流群
二维码
科研交流群
商业服务