GPTCloneBench
收藏资源简介:
GPTCloneBench是一个由萨斯喀彻温大学团队提出的语义和跨语言克隆数据集,包含37,149条单语言语义克隆对和20,770条跨语言克隆对。该数据集通过GPT模型辅助生成,并经过人工验证、功能测试和工具辅助验证。GPTCloneBench旨在解决现有数据集在规模、克隆类型和语言多样性方面的不足,适用于深度学习模型在语义克隆检测中的评估和训练,特别是在跨语言环境下的克隆检测任务。
GPTCloneBench is a semantic and cross-language clone dataset proposed by the team from the University of Saskatchewan. It contains 37,149 monolingual semantic clone pairs and 20,770 cross-language clone pairs. This dataset is generated with the assistance of GPT models, and verified via manual validation, functional testing and tool-assisted validation. GPTCloneBench aims to address the shortcomings of existing datasets in terms of scale, clone types and language diversity, and is suitable for the evaluation and training of deep learning models for semantic clone detection, especially for cross-language clone detection tasks.

- 1On the Use of Deep Learning Models for Semantic Clone Detection萨斯喀彻温大学 · 2024年



