遇见数据集

Chulalongkorn Corpus of Spoken Thai

收藏
Zenodo2026-03-05 更新2026-05-26 收录
官方服务:

资源简介:

README — Chulalongkorn Corpus of Spoken Thai (CCOST) Overview CCOST is a phonetically annotated corpus of Standard Thai comprising approximately 7 hours of spontaneous speech and 5 hours of controlled speech (word list and sentence readings) from 49 speakers. Full details are provided in the accompanying paper (see citation below). Files File Description C-COST.zip Main corpus: audio files (.wav) and annotations (.TextGrid) CCOST_metadata.csv Speaker and recording metadata (one row per file) wordlist.csv The 206-item monosyllabic word list used in the reading task sentences.txt The 25 sentences used in the reading task Interview schedule.pdf Semi-structured interview guide used for elicitation CCOST_manual.pdf Annotation manual for force alignment correction and transcription check_audio_quality_praat Praat script used for automated audio quality assessment Pittayaporn et al 2026_CCOST preprint of LREC 2026 conference proceedings (CC-BY-NC License) Metadata abbreviations Abbreviation Meaning BKK Bangkok SNR Signal-to-noise ratio dB Decibels dBFS Decibels relative to full scale s. Seconds INT Interview WL Word list ST Sentence task YOB Year of birth Year_to_BKK Year the speaker moved to Bangkok (if applicable) Age_Group Speaker age group category Audio file naming convention Files follow the format: ccost_S[n]_[task]_[filenumber] S[n]: Speaker number (e.g., S32 = 32nd speaker) Task: INT (interview), WL (word list), ST (sentence) File number: Order of recording; combined files listed as a sequence (e.g., WL_12 = sections 1 and 2 combined) License CC-BY-NC-SA 4.0. You are free to share and adapt this material for non-commercial purposes, provided you give appropriate credit and distribute any derivatives under the same license. Citation Pittayawat Pittayaporn, Cathryn Yang, Sujinat Jitwiriyanont, and James Kirby. in press. The Chulalongkorn Corpus of Spoken Thai (CCOST). In Proceedings of the 15th Language Resources and Evaluation Conference (LREC 2026), Palma de Mallorca, Spain. European Language Resources Association (ELRA).

README — 朱拉隆功泰语口语语料库(Chulalongkorn Corpus of Spoken Thai,简称CCOST) 概述 CCOST是一套针对标准泰语的语音标注语料库,包含来自49位说话者的约7小时自发口语与5小时受控口语(含单词表朗读与句子朗读任务内容)。详细信息请参阅随附论文(详见下文引用信息)。 文件 | 文件 | 描述 | | ---- | ---- | | C-COST.zip | 主语料库:包含音频文件(.wav格式)与标注文件(.TextGrid格式) | | CCOST_metadata.csv | 说话者与录音元数据(每个文件对应一行) | | wordlist.csv | 朗读任务中使用的206项单音节单词表 | | sentences.txt | 朗读任务中使用的25个句子 | | Interview schedule.pdf | 用于引导自发口语采集的半结构化访谈提纲 | | CCOST_manual.pdf | 用于强制对齐校正与转录的标注手册 | | check_audio_quality_praat | 用于自动音频质量评估的Praat脚本 | | Pittayaporn et al 2026_CCOST | LREC 2026会议论文预印本(采用CC-BY-NC许可协议) | 元数据缩写说明 | 缩写 | 含义 | | ---- | ---- | | BKK | 曼谷(Bangkok) | | SNR | 信噪比(Signal-to-noise ratio) | | dB | 分贝(Decibels) | | dBFS | 满量程分贝(Decibels relative to full scale) | | s. | 秒(Seconds) | | INT | 访谈(Interview) | | WL | 单词表任务(Word list) | | ST | 句子朗读任务(Sentence task) | | YOB | 出生年份(Year of birth) | | Year_to_BKK | 说话者移居曼谷的年份(如适用) | | Age_Group | 说话者年龄组别分类 | 音频文件命名规范 文件命名格式如下:ccost_S[n]_[task]_[filenumber] - S[n]:说话者编号(例如S32即第32位说话者) - Task:任务类型,可选INT(访谈)、WL(单词表任务)、ST(句子朗读任务) - File number:录音序号;组合文件以序列形式标注(例如WL_12代表第1与第2段内容的合并文件) 许可协议 本语料库采用CC-BY-NC-SA 4.0(知识共享署名-非商业性使用-相同方式共享4.0国际许可协议)。您可自由共享与改编本材料,但需用于非商业用途,且必须给出适当署名,衍生作品需采用相同许可协议进行分发。 引用格式 Pittayawat Pittayaporn、Cathryn Yang、Sujinat Jitwiriyanont与James Kirby. 已录用. The Chulalongkorn Corpus of Spoken Thai (CCOST). 见:第15届国际语言资源与评价会议(Language Resources and Evaluation Conference, LREC 2026)论文集,西班牙马略卡岛帕尔马。欧洲语言资源协会(European Language Resources Association, ELRA)出版。

提供机构:
Zenodo
创建时间:
2026-03-05
二维码
社区交流群
二维码
科研交流群
商业服务