eduhk-compling/GroupJ_Project
收藏资源简介:
该数据集捕捉了日本粤语学习者在音段和超音段方面面临的挑战。约一小时的语音数据来自两位社交媒体创作者(Threads上的@aa_mung和@meiceon51)。视频被下载后,音频轨道被分离,并使用Praat分割成单句片段。转录是手动完成的,并经过准确性检查,最终转换为带有对齐元数据的CSV格式。标注突出了音段错误(如发音)和超音段问题(如声调变化、停顿)。数据集在Hugging Face上公开可用,可用于语言分析、教学和NLP错误检测。
This dataset captures segmental and suprasegmental challenges faced by Japanese learners of Cantonese. Approximately one hour of learner speech was collected from two social media creators (@aa_mung and @meiceon51 on Threads). Videos were downloaded, audio tracks separated, and segmented into single‑sentence clips using Praat. Transcriptions were manually produced, checked for accuracy, and converted into CSV format with aligned metadata. Annotations highlight segmental errors (e.g., articulation) and suprasegmental issues (e.g., tone shifts, pauses). The dataset is openly available on Hugging Face and offers reuse potential for linguistic analysis, pedagogy, and NLP error detection.




