To simulate real biomedical application scenarios, we removed the reference knowledge section from the PubMedQA dataset, retaining only the Question and Final Decision as the problem and answer, formi
This deposit contains the raw per-question evaluation outputs, aggregated statistics, and figures behind every numerical claim in the manuscript "Graph-Augmented Retrieval for Biomedical Question Answ
Bio-KGR Benchmark v1.0 Bio-KGR is an automatically generated, evidence-grounded biomedical question-answering benchmark for evaluating the domain-specific reasoning abilities of large language models