遇见数据集

CN-NewsTTS Bench v0.1: A Target-Level Automatic Benchmark for Raw-Input Chinese News TTS Pronunciation

收藏
Zenodo2026-06-25 更新2026-06-28 收录
官方服务:

资源简介:

CN-NewsTTS Bench v0.1 is an open target-level automatic benchmark for evaluating raw-input Chinese news text-to-speech pronunciation accuracy. It focuses on compact written expressions that are common in Chinese news, including sports scores, ranges, military and vehicle model names, units, percentages, generation labels, brand strings, and abbreviations. The package includes the dev and public test sets, target-level annotation schema, positive and negative reading patterns, a fixed three-ASR evaluation protocol, canonical ASR transcripts, generated TTS audio artifacts, scoring scripts, seven initial commercial TTS leaderboard results, reproducibility documentation, and an arXiv preprint. The benchmark uses a Raw Input Product Track: systems are evaluated from raw text input without external text normalization, LLM rewriting, SSML, or manual text fixes. Provider-internal normalization remains part of the tested product behavior. The included audio files are generated TTS evaluation artifacts normalized to 24 kHz mono wav: 1,400 dev audio files and 5,600 public-test audio files. Code files are released under MIT. Benchmark data, documentation, ASR transcripts, target-level scores, and metadata are released under Creative Commons Attribution 4.0 International. Reuse of generated commercial TTS audio may be subject to provider/API terms and should not be treated as an unrestricted speech-training corpus without separate rights review. Release notes:Git commit packaged here: f76ff26a41b4480e6567d74df4dc9490234eb252. arXiv preprint: https://arxiv.org/abs/2606.24714. Data construction date: 2026-06-20. Initial TTS invocation date: 2026-06-22. Release audit date: 2026-06-23. Audio package contents: 7 TTS systems x 200 dev records = 1,400 dev wav files; 7 TTS systems x 800 public-test records = 5,600 public-test wav files. Provider-returned raw audio duplicates are not included; canonical audio is normalized to 24 kHz mono wav.

提供机构:
Zenodo
创建时间:
2026-06-24
二维码
社区交流群
二维码
科研交流群
商业服务