KlingTeam/UnityShotsBench
收藏资源简介:
UnityShots Benchmark 是一个多语言、多文化的 k-shot 叙事基准数据集,用于评估多镜头音频-视频生成。每个案例是一个短篇电影故事,由多个镜头讲述,具有一致的演员阵容,其身份、声音和世界必须在每个镜头切换中保持一致。该数据集包含200个故事序列(从story_001到story_200),涉及430个命名角色,每个角色都有参考身份和声音。每个故事包含3到6个镜头(平均5.0个),覆盖13种语言(包括普通话、粤语、英语、德语、西班牙语、阿拉伯语、印地语、孟加拉语、斯瓦希里语、约鲁巴语、波斯语、葡萄牙语和越南语)和6个文化区域(东亚、欧洲、南亚/东南亚、非洲、拉丁美洲、中东)。数据集支持三种输入模式:文本到视频(T2V)、图像到视频(I2V)和参考到视频(R2V),旨在公平评估跨镜头身份保持、声音/音色一致性、口型同步、场景连续性和音频-文本对齐。数据集结构包括索引文件、序列列表、参考身份肖像、参考语音剪辑和每镜头首帧锚点。数据集仅用于学术和非商业研究。
A multilingual, multi-cultural k-shot storytelling benchmark for evaluating multi-shot audio-video generation. Each case is a short cinematic story told across several shots, with a consistent cast whose identity, voice, and world must persist across every cut. The dataset includes 200 story sequences (from story_001 to story_200), featuring 430 named characters with reference identity and voice. Each story contains 3 to 6 shots (mean 5.0), covering 13 languages (Mandarin, Cantonese, English, German, Spanish, Arabic, Hindi, Bengali, Swahili, Yoruba, Persian, Portuguese, Vietnamese) and 6 cultural regions (East Asia, Europe, South/SE Asia, Africa, Latin America, Middle East). It supports three input modes: Text-to-Video (T2V), Image-to-Video (I2V), and Reference-to-Video (R2V), designed for fair evaluation of cross-shot identity preservation, voice/timbre consistency, lip-sync, scene continuity, and audio–text alignment. The dataset structure includes an index file, sequence list, reference identity portraits, reference voice clips, and per-shot first-frame anchors. It is intended for academic, non-commercial research only.




