React-to-Me Expert Evaluation Dataset
收藏资源简介:
This dataset provides a complete expert evaluation framework for benchmarking grounded versus ungrounded large language models on domain-specific biomedical question answering. It accompanies "React-to-Me: A Conversational AI System for Interactive Exploration of the Reactome Pathway Knowledgebase" and contains all materials needed to replicate a blinded comparative assessment. The evaluation set comprises 109 molecular biology questions derived from Voet & Voet, Biochemistry (2nd ed.), stratified by cognitive complexity: 76 query-like questions requiring factual recall, and 33 reasoning questions requiring synthesis and multi-step inference. For each question, the dataset includes paired responses from React-to-Me (retrieval-augmented generation grounded in Reactome pathways) and GPT-4o-mini (ungrounded baseline), plus optional textbook hint excerpt citations. Responses were evaluated using a standardized four-point ordinal rubric across three independent dimensions: factual accuracy, level of granularity (biological specificity), and relational depth (mechanistic integration). Cumulative-link mixed-effects modeling revealed that grounded responses were twice as likely to receive higher expert ratings (OR=2.01, 95% CI: 1.50-2.69, p<0.01), with performance gains generalizing across both factual and reasoning tasks. This dataset enables replication of the reported evaluation and provides a reusable benchmark for assessing domain-grounded AI systems in molecular biology. The three-metric rubric, validated through blinded expert assessment, offers a standardized framework for evaluating factual reliability, biological precision, and systems-level reasoning in biomedical question-answering systems.



