ADC
收藏资源简介:
ADC数据集由北京航空航天大学和阿里巴巴国际数字商业联合构建,旨在提升大语言模型在复杂函数调用中的鲁棒性和准确性。该数据集包含高质量的代码微调数据,并通过行级执行反馈提供细粒度的过程监督,增强模型的逻辑推理能力和函数格式遵循能力。数据来源包括CodeNet和POJ104,分别包含约1400万和5.2万条代码片段。数据集通过对抗性生成过程进一步优化,生成具有挑战性的函数调用数据,提升参数匹配的准确性。ADC数据集的应用领域主要集中在提升大语言模型在函数调用中的表现,特别是在复杂参数匹配和多样化编程场景中的能力。
The ADC dataset was jointly constructed by Beihang University and Alibaba International Digital Commerce, with the goal of improving the robustness and accuracy of large language models (LLMs) in complex function calls. This dataset includes high-quality code fine-tuning data, and provides fine-grained process supervision via line-level execution feedback, to strengthen the model's logical reasoning ability and compliance with function call formats. Its data sources are CodeNet and POJ104, which contain approximately 14 million and 52,000 code snippets respectively. The dataset is further optimized through an adversarial generation process, producing challenging function call data to enhance the accuracy of parameter matching. The application scope of the ADC dataset mainly focuses on improving the performance of LLMs in function calls, particularly their capabilities in complex parameter matching and diverse programming scenarios.

- 1ADC: Enhancing Function Calling Via Adversarial Datasets and Code Line-Level Feedback北京航空航天大学, 阿里巴巴国际数字商业 · 2024年



