marcosremar2/qwen3-omni-ptbr-qwen3tts-synthetic
收藏资源简介:
该数据集名为Qwen3 Omni PT-BR Synthetic TTS Teacher Dataset,是一个用于实验性全语音(omni speech)管道的合成葡萄牙语(巴西)语音/文本示例集合。管道流程包括:Whisper/ASR -> Tucano hidden states -> Qwen3 audio-code talker -> Qwen3-TTS tokenizer decoder。数据集内容包含50,000个合成问答/文本行(fast_pass_answer_texts.jsonl)、12,204个生成的WAV音频文件(teacher_wavs_det_v1/)、轻量级音频元数据清单(audio_metadata.jsonl)、教师清单(包括Qwen3-TTS音频代码,fast_pass_teacher_codes_det_v1.jsonl)以及数据集摘要(dataset_summary.json)。音频生成使用Qwen3-TTS教师管道,具体包括教师模型Qwen/Qwen3-TTS-12Hz-1.7B-Base、分词器/编解码器目标Qwen/Qwen3-TTS-Tokenizer-12Hz、语音引擎qwen3_tts_teacher_voice_clone、说话者IDvenus_iasini、采样率24 kHz和量化器16。目标生成50,000个示例,但此快照在12,204个音频示例处停止。数据集旨在用于训练和评估一个基于PT-BR LLM隐藏状态的紧凑Qwen3音频代码对话器,属于研究/原型用途。语音文件为Qwen3-TTS生成的合成输出,使用时需确保拥有说话者身份/参考材料的必要权利,并遵守源模型和语音参考的许可条款。
This dataset, named Qwen3 Omni PT-BR Synthetic TTS Teacher Dataset, contains synthetic Portuguese (Brazil) speech/text examples for an experimental omni speech pipeline: Whisper/ASR -> Tucano hidden states -> Qwen3 audio-code talker -> Qwen3-TTS tokenizer decoder. It includes 50,000 synthetic Q&A/text rows (fast_pass_answer_texts.jsonl), 12,204 generated WAV files (teacher_wavs_det_v1/), a lightweight audio manifest with relative WAV paths and text (audio_metadata.jsonl), a teacher manifest including Qwen3-TTS audio codes (fast_pass_teacher_codes_det_v1.jsonl), and a dataset summary (dataset_summary.json). The audio was generated using a Qwen3-TTS teacher pipeline with the teacher model Qwen/Qwen3-TTS-12Hz-1.7B-Base, tokenizer/codec target Qwen/Qwen3-TTS-Tokenizer-12Hz, voice engine qwen3_tts_teacher_voice_clone, speaker ID venus_iasini, sample rate 24 kHz, and quantizers 16. The target was 50,000 examples, but this snapshot was stopped at 12,204 generated audio examples. The dataset is intended for research/prototype use to train and evaluate a compact Qwen3-audio-code talker conditioned on PT-BR LLM hidden states. The speech files are synthetic outputs generated by Qwen3-TTS, and users must ensure they have the necessary rights for the speaker identity/reference material and comply with the licenses and terms of the source models and voice references.



