Andha-Dhun
收藏资源简介:
Andha-Dhun是由国际信息技术学院和Jio平台有限公司联合创建的首个印地语音频描述数据集,专为视障与低视力人群的媒体可访问性研究而设计。该数据集包含8部全长电影的5870条人工撰写的印地语音频描述句子,覆盖约6.3万秒的叙事内容,平均每条描述含16.1个单词,数据来源于非营利音频库AudioVault的原始音轨。其构建过程通过音频分块、语音增强、Gemini模型转录及人工校验对齐等流程,确保时间戳与视觉内容精确同步。该数据集主要应用于跨文化音视频内容可访问性研究,旨在解决印度语种音频描述资源匮乏的问题,并为自动生成系统的开发与评估提供基准。
Andha-Dhun is the first Hindi audio description dataset co-developed by the International Institute of Information Technology and Jio Platforms Limited, purpose-built for media accessibility research targeting visually impaired and low-vision individuals. This dataset includes 5,870 manually authored Hindi audio description sentences derived from 8 full-length feature films, covering approximately 63,000 seconds of narrative content, with each description averaging 16.1 words in length. The raw data is sourced from the original audio tracks of the non-profit audio repository AudioVault. Its construction pipeline encompasses audio segmentation, speech enhancement, transcription using the Gemini model, and manual verification and alignment, which ensure precise synchronization between timestamps and their associated visual content. This dataset is primarily utilized in cross-cultural audio-visual content accessibility research, with the objective of addressing the scarcity of audio description resources for Indian languages, and providing benchmarks for the development and evaluation of automatic audio description generation systems.
数据集概述
数据集名称:AndhaDhun-HindiAD
所属项目:Hindi Audio Descriptions相关工作
当前状态:即将发布(Coming soon!)
主要目标:作为首个针对印地语(Hindi)音频描述(Audio Descriptions)的研究工作而建立
数据来源:GitHub仓库(https://github.com/katha-ai/AndhaDhun-HindiAD)
备注:该数据集目前尚未正式公开,具体内容、规模和格式待后续发布



