quocanh34/data_for_synthesis
收藏资源简介:
--- dataset_info: features: - name: id dtype: string - name: sentence dtype: string - name: intent dtype: string - name: sentence_annotation dtype: string - name: entities list: - name: type dtype: string - name: filler dtype: string - name: file dtype: string - name: audio struct: - name: array sequence: float64 - name: path dtype: string - name: sampling_rate dtype: int64 - name: origin_transcription dtype: string - name: sentence_norm dtype: string - name: w2v2_large_transcription dtype: string - name: wer dtype: int64 splits: - name: train num_bytes: 3484659441 num_examples: 6729 download_size: 825836967 dataset_size: 3484659441 --- # Dataset Card for "data_for_synthesis" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集信息: 特征: - 名称: id,数据类型: 字符串 - 名称: sentence,数据类型: 字符串 - 名称: intent,数据类型: 字符串 - 名称: sentence_annotation,数据类型: 字符串 - 名称: entities,列表类型,包含子特征: - 名称: type,数据类型: 字符串 - 名称: filler,数据类型: 字符串 - 名称: file,数据类型: 字符串 - 名称: audio,结构体类型,包含子项: - 名称: array,64位浮点型音频序列 - 名称: path,数据类型: 字符串 - 名称: sampling_rate,采样率,数据类型: 64位整型 - 名称: origin_transcription,数据类型: 字符串 - 名称: sentence_norm,数据类型: 字符串 - 名称: w2v2_large_transcription,数据类型: 字符串 - 名称: wer,词错误率(Word Error Rate, WER),数据类型: 64位整型 数据集划分: - 名称: train,占用字节数: 3484659441,样本数量: 6729 下载大小: 825836967 字节 数据集总大小: 3484659441 字节 --- # 「data_for_synthesis」数据集卡片 [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集概述
数据集信息
特征
- id: 字符串类型
- sentence: 字符串类型
- intent: 字符串类型
- sentence_annotation: 字符串类型
- entities: 列表类型
- type: 字符串类型
- filler: 字符串类型
- file: 字符串类型
- audio: 结构类型
- array: 浮点数序列
- path: 字符串类型
- sampling_rate: 整数类型
- origin_transcription: 字符串类型
- sentence_norm: 字符串类型
- w2v2_large_transcription: 字符串类型
- wer: 整数类型
数据分割
- train:
- 字节数: 3484659441
- 样本数: 6729
数据集大小
- 下载大小: 825836967 字节
- 数据集大小: 3484659441 字节



