pietrorisipr-2025/corelang5-benchmark
收藏资源简介:
CoreLang5 Benchmark 是一个包含117,000个确定性计算推理问题的数据集,覆盖11个领域,包括逻辑与控制流、数学与代数、数据结构、算法、I/O与序列化、并发与异步、互操作与FFI、系统与网络、优化与分析、领域适配器以及加密与互操作。每个问题都有精确且可验证的答案,无歧义和主观性,旨在评估大型语言模型在具体计算任务上的推理能力。答案通常是数字或短字符串,可通过编程方式验证。数据集具有难度等级(从1到7)、唯一标识符、SHA-256哈希用于完整性验证,且问题提示以意大利语编写(版本5.0)。预期用途包括LLM推理评估、特定领域基准测试、逐步推理训练数据以及技术/数学问题解决的微调。
CoreLang5 Benchmark is a dataset of 117,000 deterministic computational reasoning problems across 11 domains. Each problem has an exact, verifiable answer — no ambiguity, no subjectivity. Designed for evaluating LLM reasoning on concrete computational tasks. Answers are numbers or short strings, verifiable programmatically. It features difficulty levels (1 to 7), unique identifiers, SHA-256 hashes for integrity, and Italian prompts (v5.0). Intended for LLM evaluation, domain-specific benchmarking, step-by-step reasoning training, and fine-tuning for technical/mathematical problem solving.



