遇见数据集

CodeSyntax

收藏
arXiv2022-10-26 更新2024-06-21 收录
官方服务:

资源简介:

CodeSyntax是由马里兰大学帕克分校等机构创建的大规模数据集,包含18701个程序样本,每个样本都标注了语法关系。数据集通过抽象语法树(AST)提取语法关系,主要用于评估预训练语言模型在理解代码结构方面的能力。该数据集适用于Python和Java语言,旨在解决预训练模型在代码理解任务中的性能评估问题,特别是在代码语法结构的理解上。

CodeSyntax is a large-scale dataset created by the University of Maryland, College Park and other institutions. It contains 18,701 program samples, each annotated with syntactic relationships. The dataset extracts syntactic relationships via Abstract Syntax Trees (AST), and is primarily used to evaluate the ability of pre-trained language models to understand code structures. Applicable to both Python and Java programming languages, this dataset aims to address the performance evaluation issues of pre-trained models in code understanding tasks, especially regarding the comprehension of code syntactic structures.

创建时间:
2022-10-26
二维码
社区交流群
二维码
科研交流群
商业服务