遇见数据集

HumanEval-Sci: a verification-grounded benchmark for scientific code generation

收藏
Zenodo2026-05-27 更新2026-05-29 收录
官方服务:

资源简介:

A benchmark of 73 scientific-code prompts across seven domains (physics, chemistry, biology, climate, mathematics, numerical methods, and engineering). Each prompt carries a reference solution, executable tests, and per-prompt verification targets (dimensional vectors, declared limits, and validation envelopes). The repository also includes the recorded evaluation runs behind the HumanEval-Sci pilot study and a dependency-free script that regenerates the study's supplementary tables from those records.

提供机构:
Zenodo
创建时间:
2026-05-27
二维码
社区交流群
二维码
科研交流群
商业服务