five

NoFunEval

收藏
arXiv2025-09-30 收录
下载链接:
https://aka.ms/NoFunEval
下载链接
链接失效反馈
官方服务:
资源简介:
该数据集为评估代码语言模型在功能性和非功能性需求方面的表现提供了一个基准,涵盖了代码编辑和分类任务。此外,它还包含了针对非功能性需求的定制评估指标和样本生成技术。该数据集对27个不同规模的代码语言模型进行了评估,这些模型的参数量从10亿到34亿不等,主要针对代码编辑和分类任务。

This dataset acts as a benchmark for evaluating the performance of code language models on both functional and non-functional requirements, encompassing code editing and classification tasks. Furthermore, it incorporates custom evaluation metrics and sample generation techniques specifically designed for non-functional requirements. This dataset has been used to evaluate 27 code language models of varying scales, with their parameter counts ranging from 1 billion to 3.4 billion, focusing primarily on code editing and classification tasks.
提供机构:
OpenAI
搜集汇总
数据集介绍
main_image_url
背景与挑战
背景概述
NoFunEval是一个用于评估代码语言模型在真实世界代码编辑场景中表现的数据集,特别关注超越功能正确性的需求,如运行时效率、可维护性、延迟、资源利用率和安全性。它包含多个子集,每个子集针对不同的非功能性指标,并提供了生成和评估脚本,支持模型性能的全面测试和比较。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务