Neyshekar: A Large-Scale Open Persian Speech Dataset
收藏资源简介:
Neyshekar is an open, community-driven Persian speech dataset collected via a web-based crowdsourcing platform at https://ney.shekar.io. It is designed to support research and development in text-to-speech (TTS), automatic speech recognition (ASR), speech representation learning, and other downstream Persian speech applications. The recordings are provided by a combination of volunteer contributors and paid voice actors, all of whom are native Persian speakers. Each release represents a stable snapshot of the dataset, enabling reproducible research and consistent benchmarking. CAUTION: Some clips in the v4 release (2026-05-14) have misaligned audio–transcript pairs. Using this snapshot may introduce label noise and produce unreliable training or evaluation results. Please do not use v4 and If you have already downloaded v4, please discard it and re-download v4.1.



