BMAS English
收藏资源简介:
BMAS English是一个用于二分类人类和机器文本的英语语言数据集,不仅能够识别机器生成的文本,还可以尝试确定其生成器,并针对减少检测的可检测性的对抗性攻击。数据集包含来自五个广泛应用于现实世界应用的领域的人类撰写的和人工智能生成的文本,包括reddit、新闻文章、维基百科内容、arXiv的科学摘要和通用问答。数据集旨在解决机器生成文本检测的问题,以保护真实性、确保透明度,并最大限度地减少生成式AI的潜在误用。
BMAS English is an English-language dataset for binary classification of human and machine-generated text. It not only recognizes machine-generated text but also attempts to identify its generator, and supports adversarial attacks designed to reduce the detectability of such generated content. The dataset contains human-written and AI-generated text from five domains widely used in real-world applications, including Reddit, news articles, Wikipedia content, scientific abstracts from arXiv, and general question-answering scenarios. It is intended to address the problem of machine-generated text detection, so as to safeguard textual authenticity, ensure transparency, and minimize the potential misuse of generative AI.



