遇见数据集

USPDATRO: Underrepresented Speech Dataset from Romanian language Open Data

收藏
Zenodo2023-05-05 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

USPDATRO<br> ========== Underrepresented Speech Dataset from Open Data: Case Study on the Romanian Language (USPDATRO) is a manually created Romanian language speech corpus.<br> It was created specifically using speech types that are underrepresented in other speech datasets.<br> Sources for this dataset are represented by open data available on multimedia platforms under a Creative Commons license.<br> The data was manually transcribed and aligned at segment level.<br> In addition to the text and audio files, we offer text annotations (lemmatization, part of speech tags, dependency parsing) in CoNLL-U Plus format. Each datasource is mentioned by URL in the metadata.csv file with associated license (a Creative Commons variant). Dataset structure:<br> - audio: Folder with audio segments in WAV format<br> - text: Folder with corresponding transcriptions<br> - conllup: Folder with corresponding token-based annotations<br> - metadata.csv: Contains information about each segment LICENSING This work (transcriptions, alignment, metadata, annotations) is provided under the license CC BY-NC-SA 4.0 (Attribution-NonCommercial-ShareAlike 4.0 International).<br> The license can be viewed online here: https://creativecommons.org/licenses/by-nc-sa/4.0/<br> and the full text here: https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode .<br> The original works considered for audio sources are available under their respective licenses (Creative Commons variants) as described in the metadata.csv file. <br> CONTACT Research Institute for Artificial Intelligence "Mihai Drăgănescu", Romanian Academy<br> Web: http://www.racai.ro<br> Contact emails: vasile@racai.ro

提供机构:
Zenodo
创建时间:
2023-05-05
二维码
社区交流群
二维码
科研交流群
商业服务