遇见数据集

Do Current Language Models Support Code Intelligence for R Programming Language?

收藏
Zenodo2024-10-01 更新2026-05-26 收录
官方服务:

资源简介:

This is the dataset used in the paper: Do Current Language Models Support Code Intelligence for Programming Language? This dataset contains code snippets from R programming language repositories on GitHub, paired with their corresponding natural language (NL) descriptions. It was created for research in software engineering tasks like code summarization and code search. The data was collected using the GitHub REST API and includes over 1,500 public R repositories. To ensure quality, only active, well-structured R packages with proper documentation were included. Roxygen2, a popular documentation framework, was used to extract both the code and its matching NL descriptions. The dataset is organized into three parts: base R functions (Base), functions from the tidyverse (Tidy), and a combined set (RCombine). The dataset follows the CodeSearchNet format, with a split for training, validation, and testing data, ensuring no duplicate functions.

本数据集为论文《Do Current Language Models Support Code Intelligence for Programming Language?》中所使用的研究数据集。 本数据集包含GitHub平台上R编程语言仓库的代码片段及其对应的自然语言(Natural Language,NL)描述,旨在服务于代码摘要、代码搜索等软件工程任务的相关研究。数据通过GitHub REST API采集自超过1500个公开R语言仓库。为保障数据集质量,仅纳入维护活跃、结构规范且具备完整文档的R软件包,并借助主流文档框架Roxygen2提取代码及其匹配的自然语言描述。 本数据集分为三个子集:基础R函数集(Base)、tidyverse函数集(Tidy)以及合并集(RCombine)。该数据集遵循CodeSearchNet格式,划分了训练集、验证集与测试集,且确保无重复函数。

提供机构:
Zenodo
创建时间:
2024-10-01
二维码
社区交流群
二维码
科研交流群
商业服务