Bio-KGR: Automated Evidence-Grounded Biomedical QA Benchmark Generation for LLM Reasoning Evaluation
收藏资源简介:
Bio-KGR Benchmark v1.0 Bio-KGR is an automatically generated, evidence-grounded biomedical question-answering benchmark for evaluating the domain-specific reasoning abilities of large language models (LLMs). The benchmark combines knowledge-graph-derived reasoning structures with supporting evidence retrieved from PubMed-indexed biomedical literature. Contents:- high-quality open-ended QA pairs (SynLethKG-based, after filtering)- Reasoning structure annotations (One-hop, Two-hop, Intersection, Attribute)- PubMed evidence references (PMID) for each QA pair- Generation prompts and evaluation protocols Use Cases:- LLM biomedical reasoning evaluation- Multi-hop QA benchmark- Knowledge-grounded generation evaluation- Retrieval-augmented QA system development License: CC-BY 4.0



