disco-eth/cineaudiodb
收藏资源简介:
CineAudioDB是一个用于电影音频源分离的真实世界评估集,包含真实电影和动画制作,带有真实音轨(对话、音乐、音效),用于评估电影源分离模型。与线性相加的合成数据集不同,真实制作经过非线性母带处理(如压缩、限制、侧链闪避、混响),因此发布的音轨通常不会线性相加到母带混合中。为支持两种假设下的公平评估,每个制作提供两个输入版本:线性混合(音轨线性相加)和制作混合(真实母带立体声混合)。数据集包括9个制作(7个开源动画和2个学生制作),总时长2.46小时,音频格式为FLAC(无损)、24位、48 kHz,主要为立体声(一个制作是单声道)。它还包含元数据文件(如metadata.csv和films.json),支持通过Hugging Face datasets库加载。数据集旨在隔离模型质量与混合过程不匹配,并测量在真实母带音频上的部署性能。
CineAudioDB is a real-world evaluation set for cinematic audio source separation, containing real film and animation productions with ground-truth stems (dialogue, music, sfx) for evaluating cinematic source-separation models. Unlike synthetic datasets that sum stems linearly, real productions are mixed with a non-linear mastering chain (e.g., compression, limiting, sidechain ducking, reverb), so the released stems generally do not sum to the mastered mix. To support fair evaluation under both assumptions, every production is provided in two input versions: linear_mix (additive sum of stems) and production_mix (real mastered stereo mix). The dataset includes 9 productions (7 open-source animations and 2 student productions), with a total duration of 2.46 hours, in FLAC (lossless), 24-bit, 48 kHz audio format, primarily stereo (one production is mono native). It also includes metadata files (e.g., metadata.csv and films.json) and supports loading via the Hugging Face datasets library. The dataset aims to isolate model quality from mixing-process mismatch and measure real deployment performance on professionally mastered audio.




