zomi_syllabified_human
收藏资源简介:
Zomi音节化单词数据集包含带有音节边界标记的Zomi单词。该数据集采用单一训练集划分,共包含16,609个单词样本。每个样本由两个字段组成:word字段存储原始Zomi单词,syllables字段存储用连字符标记音节边界的同一单词。该数据集专为语言学研究和自然语言处理任务设计,特别适用于音节分割、语音分析和Zomi语言处理等应用场景。数据来源于Google Sheet的自动同步,最后更新时间为2026年6月17日。
The Zomi Syllabified Words dataset contains Zomi words with syllable boundary markers. It uses a single training set split and includes 16,609 word samples. Each sample consists of two fields: the word field stores the original Zomi word, and the syllables field stores the same word with syllable boundaries marked by hyphens. The dataset is designed for linguistic research and natural language processing tasks, particularly suitable for applications such as syllable segmentation, phonetic analysis, and Zomi language processing. The data is sourced from automatic synchronization with Google Sheets, with the last update on June 17, 2026.
- 数据集名称: Zomi Syllabified Words
- 语言: ctd(Zomi语言)
- 许可证: MIT
- 规模: 10K < n < 100K(共16,614个单词)
- 任务类别: 其他
- 数据集结构: 包含一个分割(
train),两个字段:word: 原始的Zomi单词syllables: 用连字符标记音节边界的单词
- 使用示例: 可通过
load_dataset("zomi-language-corpora/zomi_syllabified_human", split="train")加载 - 来源: 自动从Google Sheets同步,最后同步时间:2026-06-17 20:02:05 UTC
- 统计数据: 总单词数:16,614




