adalat-ai/in22-legal
收藏资源简介:
IN22-Llegal是一个专门用于印度语种自动语音识别(ASR)的测试专用、分布外法律领域听写基准数据集。它包含来自IN22-Gen语料库的法律段落朗读录音,密集涵盖领域实体(如法规名称、条款编号)、正式数字(如日期、货币金额)和复杂从句结构。数据集支持三种语言:马拉雅拉姆语(ml)、印地语(hi)和卡纳达语(kn),每种语言包含80个音频样本,仅提供测试分割。每个样本包括音频、丰富转录文本(法律级目标)、说话者ID、持续时间和语言信息。数据通过母语者朗读录制,转录文本采用丰富格式(包括语法标点、格式化数字和印度文字正字法约定),并经过法律领域注释者手动验证。该数据集旨在评估在通用领域语料库上训练的ASR系统对专业法律词汇的泛化能力,作为诊断性分布外探测工具。
IN22-Legal is a test-only, out-of-distribution dictation benchmark dataset for automatic speech recognition (ASR) of Indian languages, targeting the legal domain. It contains read-aloud audio recordings of legal passages sourced from the IN22-Gen corpus, which densely covers domain-specific entities such as statutory names, clause numbers, formal numerals including dates and monetary amounts, and complex clause structures. The dataset supports three languages: Malayalam (ml), Hindi (hi), and Kannada (kn), with 80 audio samples per language, and only the test split is provided. Each sample includes audio, a rich transcription text (legal-grade reference), speaker ID, duration, and language information. The data was recorded by native speakers, and the transcriptions adopt a rich format including grammatical punctuation, formatted numerals, and Indian language orthographic conventions, and have been manually validated by legal domain annotators. This dataset is designed to evaluate the generalization capability of ASR systems trained on general-domain corpora for specialized legal vocabulary, acting as a diagnostic out-of-distribution probing tool.



