TreeBench-861: A Benchmark for Hierarchy-Sensitive Retrieval over Structured Regulatory Corpora
收藏资源简介:
TreeBench-861 is a benchmark dataset for evaluating hierarchy-sensitive retrieval over structured regulatory corpora. The benchmark contains 861 questions across five domains: tax, finance, legal, medical, and compliance. Each item is designed to test retrieval failures where semantically similar text is insufficient to identify the governing authority, especially when parent rules, child exceptions, overrides, cross-references, temporal conditions, and scope definitions change the correct answer. Each benchmark item includes gold structural annotations such as required_node_ids, distractor_node_ids, gold_path, gold_evidence, and review_status. The dataset is intended to support reproducible evaluation of retrieval-augmented generation systems, dense retrieval, BM25, hybrid retrieval, reranking, chain-of-thought prompting, LLM judge pipelines, and hierarchy-aware retrieval methods. The key benchmark finding is that high answer accuracy can mask low authority recall: non-oracle retrieval methods can reach approximately 76 percent answer accuracy while recovering only about 45 percent of required governing nodes. TreeBench-861 therefore emphasizes required-node recall as a primary metric for regulated-domain retrieval.



