cyberagent/JOR-Bench
收藏资源简介:
JOR-Bench是一个用于评估大语言模型在运筹学问题建模上的双语基准数据集,包含1,319个运筹学文字问题,提供英文和日文版本,源自五个公开的英文基准。该数据集旨在填补日文运筹学资源的空白,支持跨语言比较,以研究数学建模能力是否能在语言间迁移。数据格式为JSON Lines,每个条目包含自然语言问题描述和最优数值答案,可用于多种运筹学求解器或建模语言进行评估。
JOR-Bench is a bilingual benchmark dataset designed to evaluate the operational research problem modeling capabilities of large language models (LLMs). It contains 1,319 operational research word problems with both English and Japanese versions, derived from five publicly available English benchmark datasets. This dataset aims to fill the gap in Japanese operational research resources, and supports cross-lingual comparisons to investigate whether mathematical modeling abilities can transfer across languages. The dataset follows the JSON Lines format, where each entry includes a natural language problem description and an optimal numerical answer, and can be used for evaluation with various operational research solvers or modeling languages.




