seanghay/khmer-speech-large
收藏资源简介:
--- dataset_info: features: - name: audio dtype: audio: sampling_rate: 16000 - name: sentence dtype: string splits: - name: train num_bytes: 5686102163.1 num_examples: 19850 - name: test num_bytes: 726356614.0 num_examples: 771 download_size: 6074861609 dataset_size: 6412458777.1 --- # Dataset Card for "khmer-speech-large" [More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
## 数据集信息 特征: - 名称:音频(audio),数据类型:音频,采样率:16000 赫兹 - 名称:句子(sentence),数据类型:字符串 数据集划分: - 名称:训练集(train),占用字节数:5686102163.1,样本数量:19850 - 名称:测试集(test),占用字节数:726356614.0,样本数量:771 下载总大小:6074861609,数据集总存储大小:6412458777.1 # “高棉语语音大数据集(khmer-speech-large)”数据集卡片 [需补充更多信息](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards)
数据集概述
数据集名称
- 名称: khmer-speech-large
数据集特征
-
音频特征:
- 名称: audio
- 数据类型:
- 采样率: 16000 Hz
-
文本特征:
- 名称: sentence
- 数据类型: string
数据集分割
-
训练集:
- 示例数量: 19850
- 数据大小: 5686102163.1字节
-
测试集:
- 示例数量: 771
- 数据大小: 726356614.0字节
数据集大小
- 下载大小: 6074861609字节
- 总数据集大小: 6412458777.1字节



