遇见数据集

Data and Code Release for "Benchmarking Multilingual ASR for Tagalog and Cebuano Across Scoring Conventions and Language Coverage" (NLPIR 2026)

收藏
Zenodo2026-09-28 更新2026-10-01 收录
官方服务:

资源简介:

Release archive for the NLPIR 2026 paper "Benchmarking Multilingual ASR for Tagalog and Cebuano Across Scoring Conventions and Language Coverage" (paper NJ 105). It contains the unnormalized per-utterance hypotheses of all fourteen runs of Whisper large-v3, MMS-1B-all, and SeamlessM4T v2 Large on the FLEURS fil_ph and ceb_ph test splits; the run configurations, environment specification, and exact model revisions; the inference and scoring notebook with the four normalization conventions, scoring, alignment, and bootstrap code and its worked-example tests; all result tables; a re-execution of the published SeamlessM4T decoding path with and without special-token suppression; the language-identity audit outputs; the labels of two independent annotators for the 46-utterance validation sample; and the Cebuano–Tagalog word list used in the substitution analysis. FLEURS audio and reference transcriptions are not redistributed and should be obtained from google/fleurs (CC BY 4.0); every file joins to it on idx or utt_id. See README.md inside the archive for details.

提供机构:
Zenodo
创建时间:
2026-09-28
二维码
社区交流群
二维码
科研交流群
商业服务