PDEAgent-Bench
收藏资源简介:
PDEAgent-Bench 是一个用于评估大语言模型和AI代理在偏微分方程(PDE)求解器代码生成方面端到端能力的基准测试系统。该数据集属于文本生成任务类别,包含英语文本内容,规模小于1000个样本。作为专业领域的基准测试,它专注于PDE求解器生成任务,可用于评估AI模型在科学计算代码生成方面的性能。数据集采用CC-BY-4.0许可协议发布。
PDEAgent-Bench is a benchmark system designed to evaluate the end-to-end capabilities of Large Language Models (LLMs) and AI Agents when generating partial differential equation (PDE) solver code. This dataset falls under the category of text generation tasks, contains English textual content, and has fewer than 1000 samples in total. As a domain-specific benchmark focused on PDE solver code generation, it can be used to assess the performance of AI models in scientific computing code generation. This dataset is released under the CC-BY-4.0 license.
PDEAgent-Bench 数据集概述
基本信息
- 数据集名称:PDEAgent-Bench
- 许可证:CC-BY-4.0
- 语言:英语
- 数据集大小:少于 1,000 条样本
- 任务类别:文本生成
- 标签:基准测试、偏微分方程(PDE)
数据集描述
PDEAgent-Bench 是一个用于评估大型语言模型和 AI 智能体在端到端偏微分方程(PDE)求解器代码生成能力的基准测试系统。
引用信息
该数据集正在 NeurIPS 2026 审稿中。若在研究中引用,请参考提供的 BibTeX 格式引用。
相关链接
- GitHub 仓库:https://github.com/YusanX/pde-agent-bench




