PHYSICS
收藏资源简介:
我们引入了一个大规模、高质量且广泛具有挑战性的PHYSICS数据集,用于训练和评估,同时提供了一个Rule+Model评估框架,为增强大型模型的物理推理能力提供了新的解决方案。数据集包含16,568个样本,覆盖力学、电磁学、热力学、光学和现代物理五个领域,以及从高中到研究生四个难度级别。
We introduce a large-scale, high-quality, and broadly challenging PHYSICS dataset for training and evaluation, while also providing a Rule+Model evaluation framework that offers a novel solution for enhancing the physical reasoning capabilities of large-scale models. The dataset contains 16,568 samples, covering mechanics, electromagnetism, thermodynamics, optics, and modern physics in five domains, as well as four difficulty levels ranging from high school to graduate level.
PHYSICS数据集概述
数据集简介
- 目的:增强大型模型的物理推理能力
- 特点:大规模、高质量、广泛挑战性
- 配套框架:Rule+Model评估框架
数据规模与构成
- 总样本量:16,568个
- 训练集:14,568个(含推理路径)
- 测试集:2,000个(难度与主题平衡)
- 数据来源:100+本教科书
- 扩展方式:中英双语翻译
数据质量保证
- 模型校正
- 专家审核
领域与难度覆盖
- 五大物理领域:
- 力学
- 电磁学
- 热力学
- 光学
- 现代物理
- 四个难度级别:
- 高中及以下
- 高中竞赛级
- 非物理专业本科
- 物理专业本科/研究生
数据字段说明
- id:唯一标识符
- question:物理问题
- solution:分步解答过程
- answer:正确答案列表
- answer_type:答案类型(区间/表达式/方程/真假/多选/数值/开放)
- language:语言(中文/英文)
- domain:所属物理领域
- difficulty:难度等级
- translate:是否翻译所得
- reason_path(仅训练集):QwQ-32B生成的详细推理路径
实验评估
- 评估模型:GPT-3、Gemini-Pro-2.5、Grok-3、DeepSeek-R1等
- 评估设置:零样本
- 关键结果:
- GPT-3准确率:58.9%
- DeepSeek-R1准确率:55.3%
- 闭源与开源模型存在显著差距
- 热力学和现代物理领域挑战最大
获取方式
- 论文:https://arxiv.org/abs/2506.00022
- 数据集:https://huggingface.co/datasets/desimfj/PHYSICS
- 原始数据:https://drive.google.com/file/d/1QFGA_CTAn7_NNyBWaybvRcwdTjfrtZZF/view
引用格式
bibtex @article{zheng2025scaling, title={Scaling Physical Reasoning with the PHYSICS Dataset}, author={Zheng, Shenghe and Cheng, Qianjia and Yao, Junchi and Wu, Mengsong and Ding, Ning and Cheng, Yu and Hu, Shuyue and Bai, Lei and Zhou, Dongzhan and Cui, Ganqu and others}, journal={arXiv preprint arXiv:2506.00022}, year={2025} }




