owensong-voxcpm2-synthetic-en-v1
收藏资源简介:
# Inflect VoxCPM2 Synthetic English v1 Synthetic English speech dataset generated with [openbmb/VoxCPM2](https://huggingface.co/openbmb/VoxCPM2). ## Dataset Summary This is a multi-voice synthetic English speech dataset prepared for: - TTS fine-tuning - voice-cloning research - synthesis benchmarking - stability, robustness, and post-processing experiments The audio in this release is **synthetic**. It is not a corpus of naturally recorded human speech. ## Dataset Structure - `train/metadata.csv`: public release manifest for the `train` split - `train/audio/<voice_id>/*.wav`: audio files organized by voice id - `README.md`: dataset card and usage terms - `LICENSE_NOTICE.txt`: short release notice - `dataset_stats.json`: summary statistics for the release - `CITATION.cff`: citation metadata ## Release Statistics | Field | Value | |---|---| | Rows | 59996 | | Distinct voices | 56 | | Approx total hours | 111.81 | | Avg clip duration (s) | 6.71 | | Source model | `openbmb/VoxCPM2` | | Language | English | | HF repo | `owensong/voxcpm2-synthetic-en-v1` | Example voices: `adam_american, adam_scottish, addison_2_0, aerisita, ak, alex, amelia, amy, arabella, arthur_american, austin, belinda, bradford, charlotte, christina, chuck_miller, clara, david_american, david_australian, denzel, ...` ## Category Distribution - `conversational`: 16975 - `narrative`: 11710 - `long`: 7000 - `short`: 5319 - `descriptive`: 5310 - `emotional`: 5237 - `instructional`: 2023 - `technical`: 1800 - `question`: 1735 - `dialogue`: 1653 - `punctuation`: 1234 ## Generation Mode Distribution - `ultimate`: 59996 ## Columns | Column | Description | |---|---| | `file_name` | Relative path to the WAV file inside the split directory | | `sample_id` | Stable per-sample id derived from the filename | | `text` | Intended spoken transcript | | `voice_id` | Prompt/reference voice identifier | | `category` | Prompt/content bucket | | `generation_mode` | Generation mode used for synthesis | | `language` | Language code (`en`) | | `synthetic` | Always `true` for this release | | `source_model` | Model used to generate the synthetic speech | | `dataset_version` | Local dataset snapshot/version label | | `duration_s` | Clip duration in seconds | | `cfg_value` | CFG value used during generation when available | | `speaker_gender` | Optional prompt metadata | | `speaker_accent` | Optional prompt metadata | | `speaker_style` | Optional prompt metadata | | `control_instruction` | Optional voice-design instruction | | `prompt_text` | Optional reference prompt transcript | | `generated_at` | Generation timestamp recorded locally | ## Intended Uses - TTS and voice-cloning research - fine-tuning and adaptation experiments - benchmarking synthesis quality, stability, and robustness - speech enhancement and post-processing experiments - evaluation of expressive or long-form synthesis systems ## Out-of-Scope Uses - impersonation of real people - fraud, deception, scams, or voice phishing - surveillance or biometric speaker identification - any use that misrepresents this dataset as natural human-recorded speech ## License and Usage This dataset is released under custom terms (`license: other`). ### Permitted Use You may use this dataset for: - research - non-commercial experimentation - benchmarking and evaluation - training or fine-tuning text-to-speech and voice-cloning systems - academic, educational, and personal projects ### Attribution Requirement If you use this dataset in a paper, model, demo, application, repository, benchmark, blog post, or derivative dataset, you must provide attribution. Recommended attribution format: - **Inflect VoxCPM2 Synthetic English v1** - Hugging Face dataset page: `owensong/voxcpm2-synthetic-en-v1` - Description: synthetic English speech dataset generated with VoxCPM2 ### Prohibited Use You may not use this dataset for: - impersonation, fraud, deception, or voice phishing - generating misleading or non-consensual synthetic media - biometric identification, speaker surveillance, or identity verification of real people - unlawful, harmful, or abusive voice-cloning applications - redistributing the dataset as-is, or under a different name, without permission - representing the audio as human-recorded natural speech ### Commercial Use Commercial use is not permitted without separate written permission from the dataset publisher. ### Nature of the Data This dataset contains **synthetic English speech**, not human-recorded source speech. It is intended for TTS, voice-cloning, benchmarking, and related speech research. Users should evaluate suitability for their own use case and should not treat this dataset as a ground-truth corpus of natural human speech behavior. ## Limitations - The reference prompt audio used during generation is **not** included in this release. - Synthetic audio quality and behavior reflect the strengths and weaknesses of the upstream generator and prompt setup. - This release is best understood as a synthetic training/evaluation resource, not as a substitute for carefully licensed natural speech corpora.



