H-AIRosettaMP
收藏资源简介:
H-AIRosettaMP数据集是由博洛尼亚大学和巴黎综合理工学院的研究团队创建的,用于AI代码风格学任务。该数据集包含121,247个代码片段,涵盖10种流行的编程语言,每个片段被标记为人类编写或AI生成。数据集的创建过程基于Rosetta Code项目,通过代码翻译生成AI编写的代码片段,确保了数据集的多语言性和可重复性。该数据集主要用于检测AI生成的代码,旨在解决代码生成中的安全、知识产权和伦理问题。
The H-AIRosettaMP dataset was created by research teams from the University of Bologna and École Polytechnique for AI code stylometry tasks. This dataset contains 121,247 code snippets covering 10 popular programming languages, with each snippet labeled as either human-written or AI-generated. The dataset was developed based on the Rosetta Code project, where AI-written code snippets were generated via code translation, ensuring the multilingualism and reproducibility of the dataset. This dataset is primarily used for detecting AI-generated code, aiming to address the security, intellectual property, and ethical issues in code generation.




