FDARxBench: Benchmarking regulatory and clinical reasoning on FDA generic drug assessment
收藏资源简介:
We introduce an expert-curated, real-world benchmark for evaluating document-grounded question answering (QA) motivated by generic drug assessment, using U.S. Food and Drug Administration (FDA) drug label documents. Drug labels contain rich but heterogeneous clinical and regulatory information, making accurate question answering difficult for current language models. In collaboration with FDA regulatory assessors, we construct a multi-stage pipeline for generating high-quality, expert-curated QA examples spanning factual, multi-hop, and refusal tasks, and design evaluation protocols to assess both open-book and closed-book reasoning. Experiments across proprietary and open-weight models reveal substantial gaps in factual grounding, long-context retrieval, and safe refusal behavior.



