parambharat/mile_dataset
收藏资源简介:
IISc-MILE Tamil ASR Corpus是一个用于自动语音识别(ASR)的泰米尔语转录语音语料库。该数据集由专家生成,语言为泰米尔语,许可证为CC BY 2.0,属于单语言数据集,大小在10K到100K之间。数据集的源数据为原始数据,标签包括Tamil ASR和Speech Recognition,任务类别为automatic-speech-recognition。
IISc-MILE Tamil ASR Corpus is a Tamil transcribed speech corpus designed for automatic speech recognition (ASR). This expert-generated dataset is in Tamil, licensed under CC BY 2.0, and is a monolingual dataset with a size ranging from 10K to 100K. The source data of the dataset is raw data, with tags including Tamil ASR and Speech Recognition, and its task category is automatic-speech-recognition.
数据集概述
基本信息
- 名称: IISc-MILE Tamil ASR Corpus
- 语言: 泰米尔语 (Tamil)
- 语言创建者: 专家生成
- 许可证: CC BY 2.0
- 多语言性: 单语种
- 大小: 10K<n<100K
- 来源: 原始数据
- 标签: Tamil ASR, Speech Recognition
- 任务类别: 自动语音识别
数据集描述
- 摘要: 泰米尔语转录的语音数据集,用于自动语音识别。
- 支持的任务和排行榜: 信息待补充
- 语言: 泰米尔语
许可证信息
- 许可证: Attribution 2.0 Generic (CC BY 2.0)
引用信息
-
论文1:
@misc{mile_1, doi = {10.48550/ARXIV.2207.13331}, url = {https://arxiv.org/abs/2207.13331}, author = {A, Madhavaraj and Pilar, Bharathi and G, Ramakrishnan A}, title = {Subword Dictionary Learning and Segmentation Techniques for Automatic Speech Recognition in Tamil and Kannada}, publisher = {arXiv}, year = {2022}, }
-
论文2:
@misc{mile_2, doi = {10.48550/ARXIV.2207.13333}, url = {https://arxiv.org/abs/2207.13333}, author = {A, Madhavaraj and Pilar, Bharathi and G, Ramakrishnan A}, title = {Knowledge-driven Subword Grammar Modeling for Automatic Speech Recognition in Tamil and Kannada}, publisher = {arXiv}, year = {2022}, }
贡献者
- 贡献者: @parambharat




