遇见数据集

Arabic Audio Text Dataset Repository

收藏
Figshare2025-09-14 更新2026-04-28 收录
官方服务:

资源简介:

This dataset, from the IDEAL 2025 project, was created to evaluate how well commercial speech-to-text (STT) tools transcribe Arabic speech.Dataset ComponentsThe repository includes several key components:Audio Files: Original recordings of native Arabic speakers, categorized by length (Short, Medium, Long, Very Long).Human Transcripts: Manually created transcripts that serve as the "ground truth" for accuracy comparison.Tool Transcripts: Machine-generated transcripts from six commercial STT tools: Clipto, Maestra, Notta, Sonix, Turboscribe, and Veed.Metadata: An Excel file containing details like recording duration, topic, word counts, speaker age, gender, and transcription accuracy scores.Ethical Documentation: Includes IRB approval, ensuring data was collected ethically and personally identifiable information was removed.

创建时间:
2025-09-14
二维码
社区交流群
二维码
科研交流群
商业服务