遇见数据集

CorpusCanarioWA

收藏
Zenodo2026-05-19 更新2026-05-26 收录
官方服务:

资源简介:

This dataset is a subcorpus of WhatsApp voice-message recordings and transcriptions in Canarian Spanish (Tenerife), extracted and annotated for studies in dialectology, sociolinguistics, and spoken language processing. Content and format Segment identifier: BTFXXX_NVYYY. Time segments in format HH:MM.SSS --> HH:MM.SSS. Fields per segment: id, timestamp, whisper (automatic transcription), transcript (normalized manual transcription), sheet (source sheet/file code), transcript_raw (text as in the manual transcription, including pauses and interjections). Intended purposes and uses: Dialectal research and variation in Canarian Spanish (phonetics, prosody, lexicon, morphosyntax). Training and evaluation of automatic speech recognition and segmentation models for research purpose only. Sociolinguistic studies of colloquial usage and informal registers on messaging platforms. Annotation and quality: Manual transcriptions produced by linguists and anonymized as specified in the repository policy. Timestamps and disfluency markers (pauses, hesitations, interjections) are preserved in transcript_raw; transcript provides a normalized version for linguistic analysis. Ethical and privacy considerations: The data include colloquial language and expressions that may be sensitive or offensive.

提供机构:
Zenodo
创建时间:
2026-05-19
二维码
社区交流群
二维码
科研交流群
商业服务