CodeUltraFeedback
收藏资源简介:
CodeUltraFeedback是由蒙特利尔大学DIRO创建的一个包含10,000个复杂指令的数据集,旨在通过AI反馈调整和校准大型语言模型(LLMs)以符合编程偏好。该数据集通过14种不同的LLMs生成响应,并使用GPT-3.5作为评判标准,提供数值和文本反馈。数据集内容涵盖指令遵循、代码解释、代码复杂性和效率、代码可读性以及编码风格等五个非功能性要求(或编程偏好)。CodeUltraFeedback不仅用于校准LLMs,还支持了如UltraFeedback、AI反馈的强化学习(RLAIF)和LLM作为评判者等先进校准技术的开发。此外,该数据集还促进了CODAL-Bench的建立,这是一个评估LLMs与编程偏好对齐的基准。
CodeUltraFeedback is a dataset of 10,000 complex instructions created by the DIRO at the University of Montreal, which aims to adjust and calibrate large language models (LLMs) to align with programming preferences through AI feedback. The dataset generates responses using 14 different LLMs and provides numerical and textual feedback using GPT-3.5 as the evaluation standard. The content of the dataset covers five non-functional requirements (or programming preferences) such as instruction adherence, code explanation, code complexity and efficiency, code readability, and coding style. CodeUltraFeedback not only serves to calibrate LLMs but also supports the development of advanced calibration techniques such as UltraFeedback, AI feedback-based reinforcement learning (RLAIF), and LLMs as evaluators. Additionally, the dataset has facilitated the establishment of CODAL-Bench, a benchmark for assessing the alignment of LLMs with programming preferences.




