Quantitative Benchmarking of Large Language Models for Biochemical Pattern Recognition and Evidence-Grounded Interpretation Using Strawberry and Blackberry Response Datasets
收藏资源简介:
Large language models (LLMs) are increasingly used for scientific information processing, yet their ability to perform experimentally grounded biochemical reasoning remains insufficiently characterized. This study presents a quantitative benchmark of biochemical pattern recognition and evidence-grounded interpretation using strawberry and blackberry response datasets. A canonical benchmark comprising 220 tasks was constructed from experimentally derived biochemical ground truth and applied identically to five instruction-tuned open-weight LLMs, yielding 1,100 model–task evaluations. Tasks covered profile-level interpretation, fruit-level comparisons, class-level reasoning, numerical estimation, concentration–activity relationships, paired numerical retrieval, and categorical response-pattern recognition. Responses were evaluated using predefined task-specific criteria and classified into three mutually exclusive final states: CORRECT, INCORRECT, or REVIEW_REQUIRED. A complementary response-level forensic analysis characterized numerical grounding and unsupported quantitative content. Across the complete benchmark, 6 of 1,100 responses were classified as correct, 371 as incorrect, and 723 as review-required. Thus, 377 evaluations (34.27%) reached a definitive scoring outcome, corresponding to an observed accuracy of 0.55% across all evaluations and a scorable accuracy of 1.59% among definitively scored responses. Model-level definitive scoring coverage ranged from 20.00% to 55.91%, indicating substantial variation in response scoreability in addition to differences in correctness. Categorical pattern-recognition tasks and numerical-grounding analyses provided complementary evidence on response behavior, including incomplete, unsupported, and non-committal outputs. The benchmark demonstrates the importance of separating biochemical correctness from response scoreability when evaluating LLMs on experimentally grounded scientific tasks. The proposed framework provides a reproducible approach for quantifying biochemical reasoning performance while preserving unresolved responses as a distinct analytical outcome rather than treating them automatically as incorrect.



