MIST (Multi-region Inpainting Speech Tampering)
收藏资源简介:
MIST是由越南邮政电信技术研究院构建的大规模多语言语音修复检测数据集,覆盖6种语言(含英语、越南语等),包含59.8万条语音样本。该数据集通过LLM引导的语义替换和神经语音克隆技术生成,每条语音包含1-3个独立修复的单词片段,伪造内容仅占单条语音时长的2%-7%。数据源自Multilingual LibriSpeech和LEMAS-Dataset语料库,采用严格的跨语言语音克隆与边界优化流程生成,主要应用于音频取证领域,旨在解决多区域局部语音篡改的检测与定位难题。
MIST is a large-scale multilingual speech forgery detection dataset constructed by the Vietnam Posts and Telecommunications Institute of Technology. It covers 6 languages including English, Vietnamese and others, and contains 598,000 speech samples. Generated via LLM-guided semantic replacement and neural speech cloning technologies, each speech sample includes 1 to 3 independent forged word segments, with the forged content accounting for only 2% to 7% of the total duration of a single speech clip. The dataset is sourced from the Multilingual LibriSpeech and LEMAS-Dataset corpora, and is produced through a rigorous cross-lingual speech cloning and boundary optimization pipeline. It is primarily applied in the field of audio forensics, aiming to address the challenges of detecting and localizing multi-regional partial speech tampering.

- 1Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization邮政电信技术研究院 · 2026年



