SpeechPPL/SALMon_Spirit-LM-Expressive-normalized
收藏资源简介:
--- configs: - config_name: bg_alignment data_files: - split: train path: bg_alignment/train-* - config_name: bg_all_consistency data_files: - split: train path: bg_all_consistency/train-* - config_name: bg_domain_consistency data_files: - split: train path: bg_domain_consistency/train-* - config_name: gender_consistency data_files: - split: train path: gender_consistency/train-* - config_name: rir_consistency data_files: - split: train path: rir_consistency/train-* - config_name: sentiment_alignment data_files: - split: train path: sentiment_alignment/train-* - config_name: sentiment_consistency data_files: - split: train path: sentiment_consistency/train-* - config_name: speaker_consistency data_files: - split: train path: speaker_consistency/train-* dataset_info: - config_name: bg_alignment features: - name: task dtype: string - name: ind dtype: int64 - name: positive_audio dtype: audio - name: negative_audio dtype: audio - name: prompt_audio dtype: audio: sampling_rate: 16000 - name: continuation_audio_positive dtype: audio: sampling_rate: 16000 - name: continuation_audio_negative dtype: audio: sampling_rate: 16000 - name: negative_audio_sanity dtype: audio: sampling_rate: 16000 - name: positive_sample_tokenwise_loss sequence: float32 - name: negative_sample_tokenwise_loss sequence: float32 - name: code_frame_rate dtype: int64 - name: code_depth dtype: int64 - name: model_sampling_rate dtype: int64 - name: ppl_sanity dtype: int64 - name: positive_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: model_generated_continuation dtype: audio: sampling_rate: 16000 splits: - name: train num_bytes: 86708136 num_examples: 200 download_size: 86708136 dataset_size: 86708136 - config_name: bg_all_consistency features: - name: task dtype: string - name: ind dtype: int64 - name: positive_audio dtype: audio - name: negative_audio dtype: audio - name: audio_transition_s dtype: int64 - name: prompt_audio dtype: audio: sampling_rate: 16000 - name: continuation_audio_positive dtype: audio: sampling_rate: 16000 - name: continuation_audio_negative dtype: audio: sampling_rate: 16000 - name: negative_audio_sanity dtype: audio: sampling_rate: 16000 - name: positive_sample_tokenwise_loss sequence: float32 - name: negative_sample_tokenwise_loss sequence: float32 - name: positive_continuation_tokenwise_loss sequence: float32 - name: negative_continuation_tokenwise_loss sequence: float32 - name: prompt_sample_tokenwise_loss sequence: float32 - name: model_generated_continuation dtype: audio: sampling_rate: 16000 - name: code_frame_rate dtype: int64 - name: code_depth dtype: int64 - name: model_sampling_rate dtype: int64 - name: ppl_sanity dtype: int64 - name: positive_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: positive_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string splits: - name: train num_bytes: 222443312 num_examples: 200 download_size: 222443312 dataset_size: 222443312 - config_name: bg_domain_consistency features: - name: task dtype: string - name: ind dtype: int64 - name: positive_audio dtype: audio - name: negative_audio dtype: audio - name: audio_transition_s dtype: int64 - name: prompt_audio dtype: audio: sampling_rate: 16000 - name: continuation_audio_positive dtype: audio: sampling_rate: 16000 - name: continuation_audio_negative dtype: audio: sampling_rate: 16000 - name: negative_audio_sanity dtype: audio: sampling_rate: 16000 - name: positive_sample_tokenwise_loss sequence: float32 - name: negative_sample_tokenwise_loss sequence: float32 - name: positive_continuation_tokenwise_loss sequence: float32 - name: negative_continuation_tokenwise_loss sequence: float32 - name: prompt_sample_tokenwise_loss sequence: float32 - name: model_generated_continuation dtype: audio: sampling_rate: 16000 - name: code_frame_rate dtype: int64 - name: code_depth dtype: int64 - name: model_sampling_rate dtype: int64 - name: ppl_sanity dtype: int64 - name: positive_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: positive_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string splits: - name: train num_bytes: 226172124 num_examples: 200 download_size: 226172124 dataset_size: 226172124 - config_name: gender_consistency features: - name: task dtype: string - name: ind dtype: int64 - name: positive_audio dtype: audio - name: negative_audio dtype: audio - name: audio_transition_s dtype: int64 - name: prompt_audio dtype: audio: sampling_rate: 16000 - name: continuation_audio_positive dtype: audio: sampling_rate: 16000 - name: continuation_audio_negative dtype: audio: sampling_rate: 16000 - name: negative_audio_sanity dtype: audio: sampling_rate: 16000 - name: positive_sample_tokenwise_loss sequence: float32 - name: negative_sample_tokenwise_loss sequence: float32 - name: positive_continuation_tokenwise_loss sequence: float32 - name: negative_continuation_tokenwise_loss sequence: float32 - name: prompt_sample_tokenwise_loss sequence: float32 - name: model_generated_continuation dtype: audio: sampling_rate: 16000 - name: code_frame_rate dtype: int64 - name: code_depth dtype: int64 - name: model_sampling_rate dtype: int64 - name: ppl_sanity dtype: int64 - name: positive_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: positive_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string splits: - name: train num_bytes: 228058502 num_examples: 200 download_size: 228058502 dataset_size: 228058502 - config_name: rir_consistency features: - name: task dtype: string - name: ind dtype: int64 - name: positive_audio dtype: audio - name: negative_audio dtype: audio - name: audio_transition_s dtype: int64 - name: prompt_audio dtype: audio: sampling_rate: 16000 - name: continuation_audio_positive dtype: audio: sampling_rate: 16000 - name: continuation_audio_negative dtype: audio: sampling_rate: 16000 - name: negative_audio_sanity dtype: audio: sampling_rate: 16000 - name: positive_sample_tokenwise_loss sequence: float32 - name: negative_sample_tokenwise_loss sequence: float32 - name: positive_continuation_tokenwise_loss sequence: float32 - name: negative_continuation_tokenwise_loss sequence: float32 - name: prompt_sample_tokenwise_loss sequence: float32 - name: model_generated_continuation dtype: audio: sampling_rate: 16000 - name: code_frame_rate dtype: int64 - name: code_depth dtype: int64 - name: model_sampling_rate dtype: int64 - name: ppl_sanity dtype: int64 - name: positive_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: positive_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string splits: - name: train num_bytes: 202444443 num_examples: 200 download_size: 202444443 dataset_size: 202444443 - config_name: sentiment_alignment features: - name: task dtype: string - name: ind dtype: int64 - name: positive_audio dtype: audio - name: negative_audio dtype: audio - name: prompt_audio dtype: audio: sampling_rate: 16000 - name: continuation_audio_positive dtype: audio: sampling_rate: 16000 - name: continuation_audio_negative dtype: audio: sampling_rate: 16000 - name: negative_audio_sanity dtype: audio: sampling_rate: 16000 - name: positive_sample_tokenwise_loss sequence: float32 - name: negative_sample_tokenwise_loss sequence: float32 - name: code_frame_rate dtype: int64 - name: code_depth dtype: int64 - name: model_sampling_rate dtype: int64 - name: ppl_sanity dtype: int64 - name: positive_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: model_generated_continuation dtype: audio: sampling_rate: 16000 splits: - name: train num_bytes: 46555074 num_examples: 200 download_size: 46555074 dataset_size: 46555074 - config_name: sentiment_consistency features: - name: task dtype: string - name: ind dtype: int64 - name: positive_audio dtype: audio - name: negative_audio dtype: audio - name: audio_transition_s dtype: int64 - name: prompt_audio dtype: audio: sampling_rate: 16000 - name: continuation_audio_positive dtype: audio: sampling_rate: 16000 - name: continuation_audio_negative dtype: audio: sampling_rate: 16000 - name: negative_audio_sanity dtype: audio: sampling_rate: 16000 - name: positive_sample_tokenwise_loss sequence: float32 - name: negative_sample_tokenwise_loss sequence: float32 - name: positive_continuation_tokenwise_loss sequence: float32 - name: negative_continuation_tokenwise_loss sequence: float32 - name: prompt_sample_tokenwise_loss sequence: float32 - name: model_generated_continuation dtype: audio: sampling_rate: 16000 - name: code_frame_rate dtype: int64 - name: code_depth dtype: int64 - name: model_sampling_rate dtype: int64 - name: ppl_sanity dtype: int64 - name: positive_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: positive_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string splits: - name: train num_bytes: 223684769 num_examples: 200 download_size: 223684769 dataset_size: 223684769 - config_name: speaker_consistency features: - name: task dtype: string - name: ind dtype: int64 - name: positive_audio dtype: audio - name: negative_audio dtype: audio - name: audio_transition_s dtype: int64 - name: prompt_audio dtype: audio: sampling_rate: 16000 - name: continuation_audio_positive dtype: audio: sampling_rate: 16000 - name: continuation_audio_negative dtype: audio: sampling_rate: 16000 - name: negative_audio_sanity dtype: audio: sampling_rate: 16000 - name: positive_sample_tokenwise_loss sequence: float32 - name: negative_sample_tokenwise_loss sequence: float32 - name: positive_continuation_tokenwise_loss sequence: float32 - name: negative_continuation_tokenwise_loss sequence: float32 - name: prompt_sample_tokenwise_loss sequence: float32 - name: model_generated_continuation dtype: audio: sampling_rate: 16000 - name: code_frame_rate dtype: int64 - name: code_depth dtype: int64 - name: model_sampling_rate dtype: int64 - name: ppl_sanity dtype: int64 - name: positive_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_sample_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: positive_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string - name: negative_continuation_raw_units list: - name: hubert dtype: string - name: pitch dtype: string - name: style dtype: string splits: - name: train num_bytes: 228183961 num_examples: 200 download_size: 228183961 dataset_size: 228183961 --- # SALMon Normalized Dataset This repo preserves the SALMon per-config folder layout while normalizing mismatched schema details across model families.
# SALMon 归一化数据集(SALMon Normalized Dataset) 本仓库保留了SALMon按配置划分的文件夹布局,同时对不同模型族间存在的不匹配模式(schema)细节进行了归一化处理。 ## 配置项(configs) - 配置名称:bg_alignment(背景对齐配置) 数据文件: - 拆分方式:训练集(train),路径:`bg_alignment/train-*` - 配置名称:bg_all_consistency(全背景一致性配置) 数据文件: - 拆分方式:训练集(train),路径:`bg_all_consistency/train-*` - 配置名称:bg_domain_consistency(背景域一致性配置) 数据文件: - 拆分方式:训练集(train),路径:`bg_domain_consistency/train-*` - 配置名称:gender_consistency(性别一致性配置) 数据文件: - 拆分方式:训练集(train),路径:`gender_consistency/train-*` - 配置名称:rir_consistency(RIR一致性配置) 数据文件: - 拆分方式:训练集(train),路径:`rir_consistency/train-*` - 配置名称:sentiment_alignment(情感对齐配置) 数据文件: - 拆分方式:训练集(train),路径:`sentiment_alignment/train-*` - 配置名称:sentiment_consistency(情感一致性配置) 数据文件: - 拆分方式:训练集(train),路径:`sentiment_consistency/train-*` - 配置名称:speaker_consistency(说话人一致性配置) 数据文件: - 拆分方式:训练集(train),路径:`speaker_consistency/train-*` ## 数据集信息(dataset_info) 以下为各配置对应的数据集详情: 1. **bg_alignment(背景对齐)配置** 特征字段: - `task`:字符串(string)类型 - `ind`:64位整型(int64)类型 - `positive_audio`:音频(audio)类型 - `negative_audio`:音频(audio)类型 - `prompt_audio`:音频(audio),采样率(sampling_rate)为16000 - `continuation_audio_positive`:音频(audio),采样率(sampling_rate)为16000 - `continuation_audio_negative`:音频(audio),采样率(sampling_rate)为16000 - `negative_audio_sanity`:音频(audio),采样率(sampling_rate)为16000 - `positive_sample_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `negative_sample_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `code_frame_rate`:64位整型(int64)类型 - `code_depth`:64位整型(int64)类型 - `model_sampling_rate`:64位整型(int64)类型 - `ppl_sanity`:64位整型(int64)类型 - `positive_sample_raw_units`:列表(list),包含三个字段: - `hubert`:字符串(string)类型 - `pitch`:字符串(string)类型 - `style`:字符串(string)类型 - `negative_sample_raw_units`:列表(list),包含三个字段: - `hubert`:字符串(string)类型 - `pitch`:字符串(string)类型 - `style`:字符串(string)类型 - `model_generated_continuation`:音频(audio),采样率(sampling_rate)为16000 数据集拆分: - 训练集(train):字节数86708136,样本数200 下载大小:86708136,数据集总大小:86708136 2. **bg_all_consistency(全背景一致性)配置** 特征字段: - `task`:字符串(string)类型 - `ind`:64位整型(int64)类型 - `positive_audio`:音频(audio)类型 - `negative_audio`:音频(audio)类型 - `audio_transition_s`:64位整型(int64)类型 - `prompt_audio`:音频(audio),采样率(sampling_rate)为16000 - `continuation_audio_positive`:音频(audio),采样率(sampling_rate)为16000 - `continuation_audio_negative`:音频(audio),采样率(sampling_rate)为16000 - `negative_audio_sanity`:音频(audio),采样率(sampling_rate)为16000 - `positive_sample_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `negative_sample_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `positive_continuation_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `negative_continuation_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `prompt_sample_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `model_generated_continuation`:音频(audio),采样率(sampling_rate)为16000 - `code_frame_rate`:64位整型(int64)类型 - `code_depth`:64位整型(int64)类型 - `model_sampling_rate`:64位整型(int64)类型 - `ppl_sanity`:64位整型(int64)类型 - `positive_sample_raw_units`:列表(list),包含三个字段: - `hubert`:字符串(string)类型 - `pitch`:字符串(string)类型 - `style`:字符串(string)类型 - `negative_sample_raw_units`:列表(list),包含三个字段: - `hubert`:字符串(string)类型 - `pitch`:字符串(string)类型 - `style`:字符串(string)类型 - `positive_continuation_raw_units`:列表(list),包含三个字段: - `hubert`:字符串(string)类型 - `pitch`:字符串(string)类型 - `style`:字符串(string)类型 - `negative_continuation_raw_units`:列表(list),包含三个字段: - `hubert`:字符串(string)类型 - `pitch`:字符串(string)类型 - `style`:字符串(string)类型 数据集拆分: - 训练集(train):字节数222443312,样本数200 下载大小:222443312,数据集总大小:222443312 3. **bg_domain_consistency(背景域一致性)配置** 特征字段与`bg_all_consistency`配置一致,仅数据集拆分参数不同: 数据集拆分: - 训练集(train):字节数226172124,样本数200 下载大小:226172124,数据集总大小:226172124 4. **gender_consistency(性别一致性)配置** 特征字段与`bg_all_consistency`配置一致: 数据集拆分: - 训练集(train):字节数228058502,样本数200 下载大小:228058502,数据集总大小:228058502 5. **rir_consistency(RIR一致性)配置** 特征字段与`bg_all_consistency`配置一致: 数据集拆分: - 训练集(train):字节数202444443,样本数200 下载大小:202444443,数据集总大小:202444443 6. **sentiment_alignment(情感对齐)配置** 特征字段: - `task`:字符串(string)类型 - `ind`:64位整型(int64)类型 - `positive_audio`:音频(audio)类型 - `negative_audio`:音频(audio)类型 - `prompt_audio`:音频(audio),采样率(sampling_rate)为16000 - `continuation_audio_positive`:音频(audio),采样率(sampling_rate)为16000 - `continuation_audio_negative`:音频(audio),采样率(sampling_rate)为16000 - `negative_audio_sanity`:音频(audio),采样率(sampling_rate)为16000 - `positive_sample_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `negative_sample_tokenwise_loss`:32位浮点型(float32)序列(sequence) - `code_frame_rate`:64位整型(int64)类型 - `code_depth`:64位整型(int64)类型 - `model_sampling_rate`:64位整型(int64)类型 - `ppl_sanity`:64位整型(int64)类型 - `positive_sample_raw_units`:列表(list),包含三个字段: - `hubert`:字符串(string)类型 - `pitch`:字符串(string)类型 - `style`:字符串(string)类型 - `negative_sample_raw_units`:列表(list),包含三个字段: - `hubert`:字符串(string)类型 - `pitch`:字符串(string)类型 - `style`:字符串(string)类型 - `model_generated_continuation`:音频(audio),采样率(sampling_rate)为16000 数据集拆分: - 训练集(train):字节数46555074,样本数200 下载大小:46555074,数据集总大小:46555074 7. **sentiment_consistency(情感一致性)配置** 特征字段与`bg_all_consistency`配置一致: 数据集拆分: - 训练集(train):字节数223684769,样本数200 下载大小:223684769,数据集总大小:223684769 8. **speaker_consistency(说话人一致性)配置** 特征字段与`bg_all_consistency`配置一致: 数据集拆分: - 训练集(train):字节数228183961,样本数200 下载大小:228183961,数据集总大小:228183961




