FlatSounds
收藏资源简介:
FlatSounds是由加州大学伯克利分校、英伟达和华盛顿大学联合创建的用于视频到音频生成物理基准测试的数据集,包含185条精心录制的室内视频片段,每条时长5至10秒,聚焦于日常物体交互产生的高能量声音事件。数据集通过智能手机采集,涵盖了材料、几何形状、环境等多种物理因素的受控变化,并提供了人工标注的事件时间戳和文本描述。该数据集旨在评估生成模型对物理过程的因果理解能力,而非表面合理性,为视频到音频模型的物理正确性和时间对齐性提供了系统化的测试框架。
FlatSounds is a physics benchmark dataset for video-to-audio generation, co-created by the University of California, Berkeley, NVIDIA, and the University of Washington. It contains 185 meticulously recorded indoor video clips, each with a duration of 5 to 10 seconds, focusing on high-energy acoustic events generated by interactions between everyday objects. The dataset is collected using smartphones, covering controlled variations of multiple physical factors including materials, geometries, and ambient environments, and provides manually annotated event timestamps and textual descriptions. This dataset is designed to evaluate the causal understanding of physical processes by generative models rather than superficial plausibility, offering a systematic test framework for assessing the physical correctness and temporal alignment of video-to-audio models.
数据集概述:FlatSounds
FlatSounds 是一个用于评估视频到音频(V2A)生成模型物理推理能力的基准数据集。该研究发表于 CVPR 2026,由 NVIDIA Cosmos Lab 团队提出。
核心目标
揭示当前 V2A 模型是否真正理解声音产生的物理过程,而非仅生成“听起来合理”的音频。研究表明,现有模型更依赖文本描述而非视觉线索来推断物理属性与语义,并在此过程中牺牲了时间对齐精度。
数据集设计
- 受控反事实对:由单一物理因素变化的时间对齐视频对组成,用于隔离特定因素的声学效应。
- 案例1(Jar Fullness):玻璃罐从空到满,物理上质量增加降低共振频率(音高 F0 降低)。
- 案例2(Damping):金属棍从自由悬挂变为表面约束,物理上阻尼增加导致衰减加快、音色变暗(频谱质心降低,衰减率升高)。
- 案例3(Environment):声学环境从走廊变为软垫房间,物理上混响时间(RT60)缩短、直达声与混响声比(DRR)升高。
- 单视频模式测试:用于探测模型内部一致性与方向性趋势。
- 案例4(Piano Key Position):手在键盘上向右移动时,音高(F0)必须严格单调递增。
数据集内容
- 额外反事实对示例:
- Scraping Surface:刮擦表面从纸板变为金属,预期金属产生更明亮刺耳的声音(频谱质心、通量、滚降升高)。
- Surface Texture:表面纹理从木地板变为地毯,预期地毯产生更柔和沉闷的声音(频谱质心、滚降降低,衰减率升高)。
- Striker Material:敲击物从金属勺变为木棍,预期木棍产生更柔和的起音与较暗音色(起音时间升高,频谱质心、滚降低)。
- 生成音频样本:提供来自多个SOTA V2A模型(FoleyCrafter、Hunyuan-V2A、MMAudio-Phys、MMAudio、ThinkSound)在有/无文本描述条件下的生成结果,供用户比较。
评估指标
- 物理音频属性:
- 起音时间(Attack Time)、衰减率(Decay Rate)、音高(F0)、频谱质心(Spectral Centroid)、频谱通量(Spectral Flux)、频谱滚降(Spectral Rolloff)、时域调制(Temporal Modulation)、混响时间(RT60)、直达声与混响声比(DRR)。
- 权衡分析:可视化“时间对齐”(Temporal Alignment)与“物理-语义保真度”(Physical-Semantic Fidelity)之间的权衡。发现添加文本描述虽提升物理与语义准确度,但普遍导致时间对齐退化。
关键发现
- 模型在推理物理过程时,对文本描述的依赖性超过对视觉流的利用。
- 基于物理的评估指标与人类偏好测试具有很强的相关性。
引用
bibtex @inproceedings{li2026flatsounds, title={Benchmarking Single-Factor Physical Video-to-Audio Generation}, author={Li, Tingle and Gururani, Siddharth and Shih, Kevin and Bhatt, Gantavya and Lee, Sang-gil and Kong, Zhifeng and Goel, Arushi and Anumanchipalli, Gopala and Liu, Ming-Yu}, booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year={2026} }
数据集链接
- 主页:https://research.nvidia.com/labs/cosmos-lab/flatsounds/

- 1Benchmarking Single-Factor Physical Video-to-Audio Generation加州大学伯克利分校; 英伟达; 华盛顿大学 · 2026年




