MelmotCR dataset
收藏资源简介:
MelmotCR数据集是一个用于训练大型语言模型(LLMs)的数据集,旨在提高其在代码审查任务中的表现。该数据集由南京大学的研究团队创建,通过从GitHub Archive中收集代码审查数据,并使用最大熵原理进行知识注入和长链式思考技术进行微调。数据集包含来自开源社区的大量代码审查记录,这些记录经过筛选和预处理,以提供丰富的结构化信息,帮助LLMs分析代码审查的多个维度,包括代码功能总结、核心逻辑分析、变更影响分析和具体问题的检查。该数据集旨在解决自动化代码审查(ACR)的挑战,通过训练LLMs在上下文理解和推理方面的能力,以更好地识别代码中的潜在问题。
The MelmotCR dataset is a resource developed for training large language models (LLMs) to improve their performance on code review tasks. It was created by a research team from Nanjing University, with code review data collected from GitHub Archive, and utilizes the maximum entropy principle for knowledge injection and long-chain thinking techniques for fine-tuning. The dataset contains a large volume of code review records sourced from open-source communities, which have been filtered and preprocessed to provide rich structured information. This enables LLMs to analyze multiple dimensions of code reviews, including code function summarization, core logic analysis, change impact analysis, and specific issue detection. This dataset is designed to address the challenges of automated code review (ACR) by enhancing LLMs' contextual understanding and reasoning capabilities to better identify potential issues in code.

- 1通过南京大学,中国 · 2025年



