liepa3
收藏资源简介:
LIEPA-3(Didysis lietuvių kalbas garsynas)是一个大规模、开源的立陶宛语语音语料库,专为自动语音识别(ASR)、文本到语音(TTS)和语言学研究而构建。该数据集包含约10,000小时的语音数据,涵盖约750万个音频文件,记录了朗读、自发、语音标注和方言等多种语音类型。采集条件多样,包括录音室、录音机、广播、电视、电话和有声读物等场景,以最大化声学多样性。数据集分为四种配置:方言语音(dial)、语音标注语音(phon)、朗读语音(read)和自发语音(spon),每种配置都有详细的统计信息,如时长、说话者数量、短语数、单词数和文件大小。数据实例包含15个字段,包括音频(FLAC格式,44.1 kHz采样率)、文本转录、说话者ID、性别、年龄组、语言类型、来源、地区等。数据集创建于2025-2026年,由维陶塔斯马格努斯大学、维尔纽斯大学和立陶宛语言研究所共同创建,以知识共享署名4.0国际许可(CC BY 4.0)发布,说话者匿名处理,使用时需注意自发材料中的真实人物或事件。
LIEPA-3 (Didysis lietuvių kalbas garsynas) is a large-scale, open-source Lithuanian speech corpus designed for automatic speech recognition (ASR), text-to-speech (TTS), and linguistic research. The dataset contains approximately 10,000 hours of speech data, covering about 7.5 million audio files, which include various speech types such as read, spontaneous, phonetically annotated, and dialectal speech. It is collected under diverse conditions, including studio, dictaphone, radio, TV, telephone, and audiobook scenarios, to maximize acoustic diversity. The dataset is divided into four configurations: dialect speech (dial), phonetically annotated speech (phon), read speech (read), and spontaneous speech (spon), each with detailed statistics such as duration, number of speakers, phrases, words, and file size. Each data instance includes 15 fields, including audio (FLAC format, 44.1 kHz sample rate), text transcription, speaker ID, gender, age group, language type, source, region, etc. Created in 2025-2026 by Vytautas Magnus University (VDU), Vilnius University (VU), and the Lithuanian Language Institute (LKI), the dataset is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0), with speakers referenced anonymously, and users are advised to use spontaneous/broadcast materials responsibly due to potential mentions of real persons or events.




