遇见数据集

mims-harvard/rel

收藏
Hugging Face2026-05-25 更新2026-06-14 收录
官方服务:

资源简介:

REL是一个多学科关系推理基准数据集,用于评估大语言模型在复杂关系模式识别方面的能力。该数据集包含三个领域:代数(1,000个问题)、生物学(7,952个问题)和化学(4,016个问题)。代数任务基于Raven渐进矩阵,涉及常数行、算术级数、排列和行和等模式;生物学任务专注于检测多序列比对和系统发育树中的结构化同塑性(即趋同进化);化学任务涵盖同分异构体识别、最大公共子结构查找、同分异构体集合补全和约束基元提取等。每个任务设计用于测试模型在不同关系复杂度下的推理能力,例如识别不变性、线性趋势、图结构关系和约束满足。数据集总规模在1K到10K之间,适用于问答任务,重点关注化学和生物学领域。

REL is a multidisciplinary relational reasoning benchmark dataset designed to evaluate the capabilities of large language models (LLMs) in complex relational pattern recognition. This dataset encompasses three domains: algebra (1,000 questions), biology (7,952 questions), and chemistry (4,016 questions). Algebraic tasks are based on Raven's Progressive Matrices, involving patterns such as constant rows, arithmetic progressions, permutations, and row sums. Biological tasks focus on detecting structural homoplasy (i.e., convergent evolution) in multiple sequence alignments and phylogenetic trees. Chemical tasks cover isomer identification, maximum common substructure search, complementation of isomer sets, and constraint primitive extraction, among others. Each task is engineered to test a model's reasoning ability across varying levels of relational complexity, including invariance recognition, linear trends, graph-structured relationships, and constraint satisfaction. The total scale of the dataset falls within the range of 1K to 10K, and it is applicable to question answering tasks, with a particular emphasis on the chemistry and biology domains.

提供机构:
mims-harvard
二维码
社区交流群
二维码
科研交流群
商业服务