NewtPhys
收藏资源简介:
NewtPhys是由法国国家信息与自动化研究所等机构联合创建的4D物理标注数据集,旨在评估基础模型对牛顿力学的底层理解能力。该数据集包含730K帧高保真渲染视频,涵盖53个真实场景和109个日常物体,通过结合3D高斯泼溅技术与牛顿力学模拟,生成了包含碰撞、重力、材料属性等11种像素级物理标注的密集注释。数据集创建过程通过将真实场景的3D高斯表示转化为可模拟粒子,并扩展Simplicits模拟器以支持多物体交互和软体动力学,最终自动生成141K个视觉问答对。该数据集主要应用于计算机视觉与人工智能领域,用于系统评估视觉语言模型和视觉基础模型的物理推理能力,解决现有基准测试在视觉真实性与物理标注密度之间的权衡问题,推动下一代物理感知模型的发展。
NewtPhys is a 4D physically annotated dataset co-developed by institutions including the French National Institute for Digital Science and Technology (INRIA), designed to evaluate the fundamental understanding of Newtonian mechanics by foundation models. This dataset comprises 730K high-fidelity rendered video frames, covering 53 real-world scenarios and 109 daily objects. By integrating 3D Gaussian Splatting technology and Newtonian mechanics simulation, it generates dense annotations with 11 pixel-level physical attributes including collision, gravity, and material properties. During the dataset construction pipeline, 3D Gaussian representations of real scenes are converted into simulatable particles, and the Simplicits simulator is extended to support multi-object interaction and soft body dynamics, ultimately automatically generating 141K visual question-answer pairs. Primarily utilized in the fields of computer vision and artificial intelligence, this dataset enables systematic assessment of the physical reasoning capabilities of vision-language models and visual foundation models. It addresses the trade-off between visual realism and physical annotation density in existing benchmark tests, and facilitates the advancement of next-generation physics-aware models.
数据集概述
- 数据集名称: NewtPhys
- 核心主题: 评估基础模型对牛顿物理学的理解能力
- 数据规模: 包含 11,000 个序列,总计 730,000 帧,以 25 FPS 渲染
- 场景与物体: 使用了 53 个真实场景和 109 个物体,共 333 个物体实例(含物理属性变化)
- 标注类型: 提供 6 种密集的逐像素物理标注:
- 碰撞
- 重力
- 材料分割
- 实例分割
- 场景流
- 变形梯度
- 数据构建: 结合真实世界 DL3DV 场景与 Google Scanned Objects,通过 3D 高斯泼溅表示,并使用基于 Simplicits 的自定义牛顿求解器模拟交互
基准测试
- 视觉问答 (VQA): 包含 141,000 个 VQA 对(84,000 个多帧样本,57,000 个单帧样本),涵盖六大研究轴:材料理解、力学、空间推理、视点、时间推理和永久性
- 评估模型: 跨 54 个视觉语言模型和 10 个视觉基础模型进行系统评估
- 主要发现:
- 模型在低层牛顿物理推理方面表现薄弱,尤其在力学和材料属性估计上
- 常见的多模态常识基准与纽特物理学准确性仅存在弱相关性,表明现有常识评估不足以衡量牛顿物理学理解
- 在视觉基础模型的物理探测中,自监督模型整体表现最佳,MiDaS 是最强的全监督基线,DINO 在重力预测上表现突出
论文信息
- 论文标题: NewtPhys: Do Foundation Models Understand Newtonian Physics?
- 作者: Sebastian Cavada, Soumava Paul, Tuan-Hung Vu, Andrei Bursuc, Raoul de Charette
- 发表年份: 2026
- 会议/期刊: arXiv
- 论文地址: https://arxiv.org/abs/2606.03986
额外资源
-
数据集主页: https://astra-vision.github.io/NewtPhys
-
BibTeX 引用:
@inproceedings{cavada2026newtphys, title = {{NewtPhys}: Do Foundation Models Understand Newtonian Physics?}, author = {Sebastian Cavada and Soumava Paul and Tuan-Hung Vu and Andrei Bursuc and Raoul de Charette}, year = 2026, booktitle = {arXiv}, url = {https://arxiv.org/abs/2606.03986} }




