遇见数据集

SADA Dataset

收藏
IEEE2026-04-17 收录
官方服务:

资源简介:

This dataset contains audio recordings sourced from more than 57 TV shows provided by the Saudi Broadcasting Authority. The total number of hours published for these recordings is ~667 hours. The recordings are in Arabic, the majority are in Saudi dialects, and some are in other dialects. To enhance the usage of SADA, the dataset is split into training, validation, and testing sets. Each of validation and testing sets is around 10 hours in audio segments length while training set is 418 hours.

提供机构:
Almutairi, Raghad
二维码
社区交流群
二维码
科研交流群
商业服务