rti-bench
收藏资源简介:
RTI-Bench是首个针对印度《信息权利法案》(2005年)下中央信息委员会(CIC)决策的结构化数据集,旨在支持法律自然语言处理、公民人工智能和AI辅助司法访问的研究。数据集包含1,516个案例,分为两个来源:来源A包含1,218个经过标注的指令-响应对,来源B包含298个来自dsscic.nic.in的结构化CIC PDF决策,涵盖5位委员和3个文档格式世代(2023-2026年)。整体标签覆盖率为82.8%(1,457个主要案例),使用完全可复现的基于规则的流程提取,未使用LLM标注。数据集结构方面,来源A包含原始行索引、案例主题行、背景叙述、委员会最终指示等字段;来源B包含原始PDF文件名、文档子类型、案例编号、委员姓名等字段。数据集定义了9种结果标签,如INFORMATION_DIRECTED(委员会指示披露)和APPEAL_DISMISSED(上诉驳回),并统计了各标签分布。豁免条款分布涵盖《信息权利法案》第8(1)条下的10个子条款,共467次引用。RTI-Bench支持四个基准任务:结果预测、豁免分类、合规结果预测和通俗语言摘要。数据收集通过基于规则的方法和手动提取PDF完成,未使用LLM标注。局限性包括17.2%的主要案例带有UNKNOWN结果标签(需要人工审查),CIC PDF语料库仅涵盖5位委员,且未包含州信息委员会决策。所有数据均来自公开来源,采用CC BY 4.0许可证发布。
RTI-Bench is the first structured dataset focused on the decisions of India’s Central Information Commission (CIC) under the Right to Information (RTI) Act 2005, intended to support research in legal natural language processing, citizen AI, and AI-assisted access to justice. The dataset contains 1,516 cases divided into two sources: Source A includes 1,218 annotated instruction-response pairs, while Source B comprises 298 structured CIC judicial decisions extracted from dsscic.nic.in, covering 5 commissioners and 3 document format generations (2023–2026). The overall label coverage rate reaches 82.8% (1,457 primary cases), extracted using a fully reproducible rule-based pipeline without LLM-based annotation. Regarding dataset structure, Source A contains fields such as original row index, case subject line, background narrative, and the commission’s final directive; Source B includes fields like original PDF filename, document subtype, case number, and commissioner name. Nine outcome labels are defined, such as INFORMATION_DIRECTED (the commission orders disclosure of information) and APPEAL_DISMISSED (the appeal is dismissed), with the distribution of each label statistically summarized. Exemption clause coverage encompasses 10 subclauses under Section 8(1) of the RTI Act, with a total of 467 citations. RTI-Bench supports four benchmark tasks: outcome prediction, exemption classification, compliance outcome prediction, and plain language summarization. The dataset was collected via rule-based methods and manual PDF extraction, with no LLM-based annotation utilized. Limitations include that 17.2% of primary cases bear an UNKNOWN outcome label (requiring manual review), the CIC PDF corpus only covers 5 commissioners, and no State Information Commission decisions are included. All data is sourced from public domains and released under the CC BY 4.0 license.




