遇见数据集

google-research-datasets/great_code

收藏
Hugging Face2024-01-18 更新2024-06-15 收录
官方服务:

资源简介:

--- annotations_creators: - expert-generated language_creators: - found language: - en license: - cc-by-sa-3.0 multilinguality: - monolingual size_categories: - 1M<n<10M source_datasets: - original task_categories: - table-to-text task_ids: [] paperswithcode_id: null pretty_name: GREAT dataset_info: features: - name: id dtype: int32 - name: source_tokens sequence: string - name: has_bug dtype: bool - name: error_location dtype: int32 - name: repair_candidates sequence: string - name: bug_kind dtype: int32 - name: bug_kind_name dtype: string - name: repair_targets sequence: int32 - name: edges list: list: - name: before_index dtype: int32 - name: after_index dtype: int32 - name: edge_type dtype: int32 - name: edge_type_name dtype: string - name: provenances list: - name: datasetProvenance struct: - name: datasetName dtype: string - name: filepath dtype: string - name: license dtype: string - name: note dtype: string splits: - name: train num_bytes: 14705534822 num_examples: 1798742 - name: validation num_bytes: 1502956919 num_examples: 185656 - name: test num_bytes: 7880762248 num_examples: 968592 download_size: 23310374002 dataset_size: 24089253989 --- # Dataset Card for GREAT ## Table of Contents - [Dataset Description](#dataset-description) - [Dataset Summary](#dataset-summary) - [Supported Tasks and Leaderboards](#supported-tasks-and-leaderboards) - [Languages](#languages) - [Dataset Structure](#dataset-structure) - [Data Instances](#data-instances) - [Data Fields](#data-fields) - [Data Splits](#data-splits) - [Dataset Creation](#dataset-creation) - [Curation Rationale](#curation-rationale) - [Source Data](#source-data) - [Annotations](#annotations) - [Personal and Sensitive Information](#personal-and-sensitive-information) - [Considerations for Using the Data](#considerations-for-using-the-data) - [Social Impact of Dataset](#social-impact-of-dataset) - [Discussion of Biases](#discussion-of-biases) - [Other Known Limitations](#other-known-limitations) - [Additional Information](#additional-information) - [Dataset Curators](#dataset-curators) - [Licensing Information](#licensing-information) - [Citation Information](#citation-information) - [Contributions](#contributions) ## Dataset Description - **Homepage:** None - **Repository:** https://github.com/google-research-datasets/great - **Paper:** https://openreview.net/forum?id=B1lnbRNtwr - **Leaderboard:** [More Information Needed] - **Point of Contact:** [More Information Needed] ### Dataset Summary [More Information Needed] ### Supported Tasks and Leaderboards [More Information Needed] ### Languages [More Information Needed] ## Dataset Structure ### Data Instances Here are some examples of questions and facts: ### Data Fields [More Information Needed] ### Data Splits [More Information Needed] ## Dataset Creation ### Curation Rationale [More Information Needed] ### Source Data [More Information Needed] #### Initial Data Collection and Normalization [More Information Needed] #### Who are the source language producers? [More Information Needed] ### Annotations [More Information Needed] #### Annotation process [More Information Needed] #### Who are the annotators? [More Information Needed] ### Personal and Sensitive Information [More Information Needed] ## Considerations for Using the Data ### Social Impact of Dataset [More Information Needed] ### Discussion of Biases [More Information Needed] ### Other Known Limitations [More Information Needed] ## Additional Information ### Dataset Curators [More Information Needed] ### Licensing Information [More Information Needed] ### Citation Information [More Information Needed] ### Contributions Thanks to [@abhishekkrthakur](https://github.com/abhishekkrthakur) for adding this dataset.

原始信息汇总

数据集描述

数据集概述

  • annotations_creators: expert-generated
  • language_creators: found
  • language: en
  • license: cc-by-sa-3.0
  • multilinguality: monolingual
  • size_categories: 1M<n<10M
  • source_datasets: original
  • task_categories: table-to-text

数据集结构

数据字段

  • id: int32
  • source_tokens: sequence: string
  • has_bug: bool
  • error_location: int32
  • repair_candidates: sequence: string
  • bug_kind: int32
  • bug_kind_name: string
  • repair_targets: sequence: int32
  • edges: list:
    • before_index: int32
    • after_index: int32
    • edge_type: int32
    • edge_type_name: string
  • provenances: list:
    • datasetProvenance: struct:
      • datasetName: string
      • filepath: string
      • license: string
      • note: string

数据分割

  • train:
    • num_bytes: 14705534822
    • num_examples: 1798742
  • validation:
    • num_bytes: 1502956919
    • num_examples: 185656
  • test:
    • num_bytes: 7880762248
    • num_examples: 968592

数据集大小

  • download_size: 23310374002
  • dataset_size: 24089253989
搜集汇总
数据集介绍
google-research-datasets/great_code 数据集图片
构建方式
在程序语言处理与软件工程领域,精准定位代码缺陷并推荐修复方案是极具挑战性的任务。GREAT数据集正是为应对这一难题而构建,其数据源自对开源代码仓库的深度挖掘与专家标注。构建过程中,首先从海量源代码中提取出具有明确错误位置与修复候选的代码片段,随后由领域专家对每个实例进行人工审核,标注出是否存在缺陷(has_bug)、错误所在索引(error_location)、修复目标(repair_targets)以及缺陷类型(bug_kind)。此外,数据集还引入了图结构信息(edges),通过记录代码元素间的先后关系与类型,为模型提供更丰富的上下文依赖。最终形成了包含约180万训练样本、18.5万验证样本及96.8万测试样本的大规模语料库。
特点
GREAT数据集的核心特点在于其多维度的精细化标注与结构化设计。每个样本不仅包含原始的源代码令牌序列(source_tokens),还附带了布尔型的缺陷标记、整数型的错误位置索引以及字符串型的修复候选列表,使得模型能够同时学习缺陷检测与自动修复两项任务。尤为突出的是,数据集创新性地引入了图边信息(edges),以before_index、after_index和edge_type三元组形式编码代码元素间的拓扑关系,这为基于图神经网络的程序理解模型提供了天然的结构化输入。此外,缺陷类型(bug_kind)与来源信息(provenances)的标注,进一步增强了数据集的实用性与可解释性。整体数据规模超过240亿字节,覆盖了丰富的代码模式与错误场景。
使用方法
使用GREAT数据集时,研究人员可将其加载为标准的表格到文本(table-to-text)任务格式,通过HuggingFace Datasets库便捷地访问各划分数据。对于缺陷检测任务,可直接利用has_bug字段作为二分类标签,结合source_tokens序列设计序列分类或Transformer模型。对于修复任务,则可依据error_location与repair_candidates构建序列到序列的生成模型,或利用repair_targets作为多标签分类目标。图结构信息edges可被解析为邻接矩阵或边列表,用于训练图卷积网络(GCN)或图注意力网络(GAT)。数据集提供了完整的训练、验证与测试划分,建议在模型评估时统一采用官方测试集以保持结果可比性。
背景与挑战
背景概述
在软件工程领域,代码缺陷的自动检测与修复一直是研究的热点与难点。由Google Research团队于2019年创建的GREAT(Graph-based Reasoning for Error-localization and ATtribution)数据集,旨在为基于图的程序错误定位与修复提供大规模监督学习基准。该数据集包含超过280万个代码片段,每个样本均标注了是否存在缺陷、错误位置及修复候选,并提供了丰富的语法与语义边信息。GREAT的提出推动了代码表示学习与程序修复技术的进步,尤其在利用图神经网络进行缺陷推理方面产生了深远影响,成为该方向的重要评估平台。
当前挑战
GREAT数据集所面临的挑战主要体现在两个方面。其一,在领域问题层面,代码缺陷检测与定位任务本身具有高度复杂性,缺陷类型多样且分布不均衡,同时代码的语义等价变换可能导致错误表象的千差万别,使得模型难以泛化。其二,在构建过程中,大规模代码数据的收集与清洗面临许可证兼容性、代码质量参差不齐等问题;人工标注错误位置与修复方案需要深厚的编程知识,成本高昂且一致性难以保证;此外,图结构的构建与边类型的定义亦需精心设计,以确保能够有效捕获代码的语法与语义依赖关系。
常用场景
经典使用场景
GREAT(Graph-based Representations for Error Analysis and Translation)数据集是软件工程与自然语言处理交叉领域的典范性资源,专门用于程序代码中的错误定位与修复任务。其最经典的使用场景是作为图神经网络(GNN)模型的训练与评估基准,研究者通过解析代码的语法与语义图结构,结合数据集提供的细粒度错误位置与修复候选信息,训练模型自动识别代码片段中的缺陷并生成修复方案。该数据集涵盖超过290万条样本,每条样本包含源代码令牌、错误标记、修复候选以及丰富的图边信息,为从代码结构层面理解程序错误提供了坚实的实验平台。
实际应用
在实际工业应用中,GREAT数据集驱动的模型被广泛集成于智能编程助手与持续集成流水线中,用于实时检测代码提交中的潜在缺陷。例如,在大型软件仓库的代码审查环节,基于该数据集训练的图神经网络能够精准定位错误行并推荐修复方案,大幅减少人工排查时间。此外,该数据集还支持教育场景中的编程练习自动批改,通过比对学习者代码与标准修复路径,提供个性化错误反馈。其结构化标注格式亦便于迁移至嵌入式系统、金融交易平台等领域的代码质量保障工具中。
衍生相关工作
GREAT数据集催生了多项具有深远影响的经典工作。以GraphCodeBERT为代表的预训练模型,借鉴了其图边信息编码思想,将代码的抽象语法树与数据流图融入语言模型预训练,刷新了多项代码智能任务基准。另一里程碑式工作是CoCoNuT,利用该数据集训练基于图卷积网络的修复系统,首次在大型编程竞赛数据集上实现超越传统模板方法的修复准确率。此外,基于GREAT的图注意力网络研究揭示了错误传播路径的可解释性机制,推动了可解释人工智能在软件工程领域的应用浪潮。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务