CyberHumanAI
收藏资源简介:
CyberHumanAI数据集由阿拉伯美国大学等机构创建,旨在通过机器学习技术检测教育内容中的AI生成文本,以维护学术诚信。该数据集包含1000条网络安全相关的段落,其中500条由人类撰写,500条由ChatGPT生成。数据来源于维基百科API,经过预处理后用于训练和测试机器学习模型。数据集的应用领域主要集中在教育领域,帮助识别学生作业中的AI生成内容,确保学术诚信。通过该数据集,研究者可以开发出高效的AI文本检测工具,提升教育环境中的透明度和公平性。
The CyberHumanAI dataset was developed by institutions including Arab American University, aiming to detect AI-generated text in educational content via machine learning techniques to uphold academic integrity. It contains 1000 cybersecurity-related paragraphs, among which 500 were written by humans and 500 were generated by ChatGPT. The data is sourced from the Wikipedia API, and after preprocessing, it is used for training and testing machine learning models. The dataset is mainly applied in the education sector, helping identify AI-generated content in student assignments to ensure academic integrity. With this dataset, researchers can develop efficient AI text detection tools, thereby enhancing transparency and fairness in educational environments.

- 1Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity阿拉伯美国大学, 哥伦比亚大学, 东密歇根大学, 德克萨斯农工大学 · 2025年



