遇见数据集

CORE-LLM-Bench: A Controlled Neurosymbolic Benchmark for Ontology-Grounded Reasoning in Large Language Models

收藏
Zenodo2026-09-25 更新2026-10-01 收录
官方服务:

资源简介:

CORE-LLM-Bench is a controlled neurosymbolic benchmark for evaluating ontology-grounded reasoning in large language models. Version 1.1.0 contains 9,048 unique question-hop instances derived from four ontology datasets: Family, Pizza 100, Pizza 250, and OWL2Bench. The benchmark contains 6,032 binary question-answering (BQA) instances and 3,016 open-ended question-answering (OEQA) instances. Benchmark instances support 1-hop and 2-hop ontology contexts and aligned natural-language (NL), formal-symbolic (FS), and entity-abstracted (AR) representations. Ground-truth answers and explanation structures are derived using symbolic ontology reasoning, enabling controlled evaluation according to task formulation, representation, semantic abstraction, context depth, explanation-based reasoning complexity, and reasoning type. The release includes benchmark data, reasoning metadata, provenance and licensing documentation, release manifests and checksums, and scripts supporting benchmark generation, validation, Hugging Face export, and reproducibility. The canonical source-code release is CORE-LLM-Bench v1.1.0 on GitHub. Licensing is component-specific. Repository software is MIT licensed; separable original CORE-LLM-Bench benchmark content is CC BY 4.0; Pizza-derived material follows CC BY 3.0; OWL2Bench-derived material follows Apache-2.0; and Family/FHKB-derived material follows CC BY-SA 3.0. See NOTICE.md for detailed provenance, attribution, and modification information.

提供机构:
Zenodo
创建时间:
2026-09-25
二维码
社区交流群
二维码
科研交流群
商业服务