Pratik/Gujarati_OpenSLR
收藏资源简介:
OpenSLR is a site devoted to hosting speech and language resources, such as training corpora for speech recognition, and software related to speech recognition. They intend to be a convenient place for anyone to put resources that they have created, so that they can be downloaded publicly. They aim to provide a central, hassle-free place for others to put their speech resources. see there http://www.openslr.org/contributions.html #Supported Task Automatic Speech Recognition #Languages Gujarati Identifier: SLR78 Summary: Data set which contains recordings of native speakers of Gujarati. Category: Speech License: Attribution-ShareAlike 4.0 International Downloads (use a mirror closer to you): about.html [1.5K] (Information about the data set ) Mirrors: [China] LICENSE [20K] (License information for the data set ) Mirrors: [China] line_index_female.tsv [423K] (Lines recorded by the female speakers ) Mirrors: [China] line_index_male.tsv [393K] (Lines recorded by the male speakers ) Mirrors: [China] gu_in_female.zip [917M] (Archive containing recordings from female speakers ) Mirrors: [China] gu_in_male.zip [825M] (Archive file recordings from male speakers ) Mirrors: [China] About this resource: This data set contains transcribed high-quality audio of Gujarati sentences recorded by volunteers. The data set consists of wave files, and a TSV file (line_index.tsv). The file line_index.tsv contains a anonymized FileID and the transcription of audio in the file. The data set has been manually quality checked, but there might still be errors. Please report any issues in the following issue tracker on GitHub. https://github.com/googlei18n/language-resources/issues See LICENSE file for license information. Copyright 2018, 2019 Google, Inc. If you use this data in publications, please cite it as follows: @inproceedings{he-etal-2020-open, title = {{Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems}}, author = {He, Fei and Chu, Shan-Hui Cathy and Kjartansson, Oddur and Rivera, Clara and Katanova, Anna and Gutkin, Alexander and Demirsahin, Isin and Johny, Cibu and Jansche, Martin and Sarin, Supheakmungkol and Pipatsrisawat, Knot}, booktitle = {Proceedings of The 12th Language Resources and Evaluation Conference (LREC)}, month = may, year = {2020}, address = {Marseille, France}, publisher = {European Language Resources Association (ELRA)}, pages = {6494--6503}, url = {https://www.aclweb.org/anthology/2020.lrec-1.800}, ISBN = "{979-10-95546-34-4}, }
OpenSLR是一个致力于托管语音与语言资源的平台,涵盖语音识别训练语料库以及相关语音识别软件。该平台旨在为所有用户提供便捷的资源上传渠道,令创作者发布的资源可被公开下载。 其目标是为用户搭建一个集中化、无冗余流程的语音资源上传平台,相关贡献指引可参见:http://www.openslr.org/contributions.html # 支持任务:自动语音识别(Automatic Speech Recognition) # 支持语言:古吉拉特语(Gujarati) 数据集标识符:SLR78 数据集摘要:本数据集收录古吉拉特语母语者的语音录制内容。 数据集类别:语音 许可协议:署名-相同方式共享4.0国际版(Attribution-ShareAlike 4.0 International) 下载项(请选择距离您较近的镜像站点): about.html [1.5K] (数据集说明文件) 镜像:[中国] LICENSE [20K] (数据集许可信息文件) 镜像:[中国] line_index_female.tsv [423K] (女性录制者对应的语音行索引文件) 镜像:[中国] line_index_male.tsv [393K] (男性录制者对应的语音行索引文件) 镜像:[中国] gu_in_female.zip [917M] (包含女性录制者语音的压缩包) 镜像:[中国] gu_in_male.zip [825M] (包含男性录制者语音的压缩包) 镜像:[中国] 本资源说明: 本数据集包含志愿者录制的古吉拉特语语句的高质量转录语音。数据集由波形音频文件与TSV格式索引文件line_index.tsv构成。line_index.tsv文件包含匿名化处理后的文件ID以及对应音频的转录文本。 本数据集已通过人工质量校验,但仍可能存在疏漏。 若发现任何问题,请前往以下GitHub问题追踪页面提交反馈:https://github.com/googlei18n/language-resources/issues 许可条款详见LICENSE文件。 版权所有 © 2018、2019 谷歌公司(Google, Inc.) 若您在学术出版物中使用本数据集,请按以下格式引用: @inproceedings{he-etal-2020-open, title = {{Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems}}, author = {He, Fei and Chu, Shan-Hui Cathy and Kjartansson, Oddur and Rivera, Clara and Katanova, Anna and Gutkin, Alexander and Demirsahin, Isin and Johny, Cibu and Jansche, Martin and Sarin, Supheakmungkol and Pipatsrisawat, Knot}, booktitle = {Proceedings of The 12th Language Resources and Evaluation Conference (LREC)}, month = may, year = {2020}, address = {Marseille, France}, publisher = {European Language Resources Association (ELRA)}, pages = {6494--6503}, url = {https://www.aclweb.org/anthology/2020.lrec-1.800}, ISBN = "{979-10-95546-34-4}, }
数据集概述
基本信息
- 数据集标识符:SLR78
- 类别:Speech
- 支持任务:Automatic Speech Recognition
- 语言:Gujarati
内容描述
- 数据集内容:包含Gujarati语的录音,由志愿者录制,包含高质量的音频文件和TSV格式文件(line_index.tsv)。
- 文件详情:
line_index_female.tsv:女性发言者的录音索引,大小423K。line_index_male.tsv:男性发言者的录音索引,大小393K。gu_in_female.zip:女性发言者的录音档案,大小917M。gu_in_male.zip:男性发言者的录音档案,大小825M。
版权与许可
- 许可证:Attribution-ShareAlike 4.0 International
- 版权声明:Copyright 2018, 2019 Google, Inc.
引用信息
-
引用格式:
@inproceedings{he-etal-2020-open, title = {Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems}, author = {He, Fei and Chu, Shan-Hui Cathy and Kjartansson, Oddur and Rivera, Clara and Katanova, Anna and Gutkin, Alexander and Demirsahin, Isin and Johny, Cibu and Jansche, Martin and Sarin, Supheakmungkol and Pipatsrisawat, Knot}, booktitle = {Proceedings of The 12th Language Resources and Evaluation Conference (LREC)}, month = may, year = {2020}, address = {Marseille, France}, publisher = {European Language Resources Association (ELRA)}, pages = {6494--6503}, url = {https://www.aclweb.org/anthology/2020.lrec-1.800}, ISBN = "{979-10-95546-34-4}, }




