SVGSA24
收藏资源简介:
SVGSA24数据集是由索尼AI和索尼集团创建的音频-视觉数据集,专门用于空间对齐的音频-视频生成研究。该数据集包含5031个视频片段,每个视频时长为5秒,分辨率为256×256,帧率为4fps,音频为立体声,采样率为16 kHz。数据集的内容主要来源于STARSS23数据集,经过转换和筛选,保留了屏幕上的语音和乐器声音事件。创建过程中,数据集通过固定视角将全景视频和FOA音频转换为透视视频和立体声音频,并进行了精细的筛选和处理。SVGSA24数据集主要应用于虚拟现实、世界模拟、机器人感知和导航等领域,旨在解决音频与视频在空间上的对齐问题,提升沉浸式体验的真实感。
The SVGSA24 dataset is an audio-visual dataset developed by Sony AI and Sony Group, specifically tailored for research on spatially aligned audio-visual generation. It comprises 5031 video clips, each with a duration of 5 seconds, a resolution of 256×256, a frame rate of 4 fps, stereo audio tracks, and a sampling rate of 16 kHz. The majority of its content originates from the STARSS23 dataset, which underwent conversion and screening procedures to retain only speech and musical instrument sound events displayed on-screen. During its development, panoramic videos and FOA audio were converted into perspective videos and stereo audio via a fixed viewpoint, followed by meticulous screening and processing. The SVGSA24 dataset is primarily utilized in domains including virtual reality, world simulation, robot perception and navigation, with the goal of addressing the spatial alignment issue between audio and video and enhancing the realism of immersive experiences.

- 1SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation索尼AI, 索尼集团 · 2024年



