官方服务:
资源简介:
81.7h Kusaal ASR - 30820 clips, 16kHz, book splits
应用场景:
创建时间:
2026-05-23
相关数据集
Exqrch/IndonesianNMT
该数据集用于论文《Replicable Benchmarking of Neural Machine Translation (NMT) on Low-Resource Local Languages in Indonesia》,包含两种类型的数据:单语(*.txt)和双语(*.tsv)。数据集涉及的语言包括印尼语(id)、爪哇语(jv)、巽他语(su)、巴厘语(ban)和米南加保语(min)。使
Hugging Face2024-01-22 更新140
LORELEI Tagalog Representative Language Pack
Introduction LORELEI Tagalog Representative Language Pack consists of Tagalog monolingual text, Tagalog-English parallel text, annotations, supplemental resources and related software tool
DataCite Commons2025-05-06 更新80
ltg/normistral-fluency-annotation
--- license: apache-2.0 --- Manual fluency annotations for [Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages](https://arxiv.org/abs/2512.08777) ### Citation ```bib
Hugging Face2025-12-10 更新40
KenSwQuAD – A Question Answering Dataset for Swahili Low Resource Language
This research developed a Kencorpus Swahili Question Answering Dataset KenSwQuAD from raw data of Swahili language, which is a low resource language predominantly spoken in Eastern African and also ha
DataONE2023-11-21 更新80
jamalimubashirali/sindhi-pretraining-corpus-part3
--- dataset_info: features: - name: book_id dtype: int64 - name: url dtype: string - name: source dtype: string - name: text dtype: string - name: chunk_index dtype: in
Hugging Face2026-03-23 更新60



