遇见数据集

hebrew-exp

收藏
魔搭社区2026-04-28 更新2026-07-15 收录
官方服务:

资源简介:

#### Why creating this dataset? We present a high-quality Hebrew speech transcription dataset generated using the **Whisper Turbo** model. In contrast, **HebDB** (https://pages.cs.huji.ac.il/adiyoss-lab/HebDB/) relies on **Whisper Large**, which is demonstrably inferior to Whisper Turbo in transcription accuracy and robustness. Furthermore, HebDB distributes audio files with **Hebrew filenames**, an avoidable design choice that introduces unnecessary friction into modern training pipelines and preprocessing workflows. Our dataset is **deliberately engineered for research and large-scale training**, requiring no additional normalization or restructuring, and is immediately usable for training a wide range of speech and language models.

提供机构:
maas
创建时间:
2026-01-22
二维码
社区交流群
二维码
科研交流群
商业服务