遇见数据集

SnipGen: A Mining Repository Framework for Evaluating LLMs for Code

收藏
Zenodo2025-01-10 更新2026-05-26 收录
官方服务:

资源简介:

This a repostiroy that contains datasets for SnipGEN replication. Each testbed is contained on a JSON file. There are three JSON file with curated data and tuned for a SE task. The tar.gz contains the raw data collected after mining github repostiries. The first testbed is the summarization taks curated. This file contains the name of the repostory the snippet comes from, the file name, the commit message, the snippet and the linked documentation. This testbed is used for completing code from code description and the combination of docstring and code. The code completion file, contains the used prompts for control, Treatment1 (T1) and Treatment(T2) and the predicted outcome from chatGPT.

本仓库(repository)用于存储SnipGEN复现所需的数据集。所有测试基准(testbed)均以JSON文件形式存储,其中三份JSON文件经过数据精选,并针对软件工程(Software Engineering,SE)任务完成调优。该tar.gz压缩包内含从GitHub仓库挖掘得到的原始数据集。 首个测试基准面向精选摘要任务构建。该文件包含代码片段来源仓库名称、文件名、提交说明(commit message)、代码片段以及关联的文档内容。此测试基准可用于基于代码描述完成代码补全,以及结合文档字符串与代码开展相关任务。 代码补全文件中包含对照组、处理组1(T1)与处理组2(T2)所使用的提示词(prompt),以及ChatGPT生成的预测结果。

提供机构:
Zenodo
创建时间:
2024-12-05
二维码
社区交流群
二维码
科研交流群
商业服务