遇见数据集

bepipeV/Vinayak-Multistep-Recursive-Reasoning-Benchmark

收藏
Hugging Face2026-05-15 更新2026-05-31 收录
官方服务:

资源简介:

Vinayak多步递归推理基准(VMRRB)是一个大规模提示基准,旨在评估前沿AI系统的高级推理、递归依赖解析、加密任务遍历和鲁棒性能力。该基准测试模型的能力包括:执行递归多步推理、解析相互依赖的问题链、执行加密依赖遍历、使用计算答案进行链式解密、解析噪声数学表达式、在极长上下文中保持一致性、处理递归执行工作流以及在大型上下文输入中保持记忆。数据集包含约1000万个相互依赖的推理任务,以加密形式呈现,问题需按依赖顺序解决,中间答案作为后续解密密钥。基准设计为连续评估提示,强调全上下文加载,不支持任意分块。数据集仅用于评估和基准测试,不提供训练分割。

The Vinayak Multistep Recursive Reasoning Benchmark (VMRRB) is a large-scale prompt-based benchmark designed to evaluate advanced reasoning, recursive dependency resolution, encrypted task traversal, and robustness capabilities of frontier AI systems. The benchmark evaluates a models ability to perform recursive multistep reasoning, resolve interdependent question chains, execute encrypted dependency traversal, perform chained decryption using computed answers, parse noisy mathematical expressions, maintain consistency across extremely long contexts, handle recursive execution workflows, and preserve memory across large contextual inputs. It contains approximately 10 million interdependent reasoning tasks, where questions are encrypted, dependencies must be recursively resolved, intermediate answers become future decryption keys, and noise must be filtered semantically. The benchmark operates as a continuous evaluation prompt, intended exclusively for evaluation and benchmarking, with no training split provided.

提供机构:
bepipeV
二维码
社区交流群
二维码
科研交流群
商业服务