Lumina-Math-Foundations-1B
收藏资源简介:
Lumina-Math-Foundations-1B 是一个工业级规模的基础数学推理数据集,语言为印尼语。该数据集基于V18.0自适应面向对象编程(OOP)与加密熵引擎构建,旨在支持数学推理、思维链(CoT)以及智能体(agentic)微调等任务。数据集包含三个核心文本字段:problem(问题)、step_solving(逐步求解过程)和answer(答案),内容均为纯自然语言与LaTeX数学表达式,无硬编码提示词。此外,数据集还提供自适应代码架构字段:code(代码)和tool_calls_json(工具调用JSON),其中高难度问题强制使用企业级@dataclass/OOP类,而简单问题则根据操作系统随机熵(os.urandom)动态平衡函数片段与类实现。解耦的工具元数据字段包括tool_calls_json、tool_responses_json和verification_payload,作为独立的JSON字符串,可用于智能体微调和基于程序化奖励的RLVR验证。数据规模约为100M至1B个样本,适用于文本生成、问答、数学推理、代码生成等场景。
Lumina-Math-Foundations-1B is an industrial-scale foundational mathematical reasoning dataset, with language in Indonesian. The dataset is built based on the V18.0 adaptive object-oriented programming (OOP) and encrypted entropy engine, designed to support tasks such as mathematical reasoning, chain-of-thought (CoT), and agentic fine-tuning. The dataset contains three core text fields: problem, step_solving, and answer, all in pure natural language and LaTeX mathematical expressions, without hardcoded prompts. Additionally, the dataset provides adaptive code architecture fields: code and tool_calls_json, where difficult problems enforce enterprise-level @dataclass/OOP classes, while simple problems dynamically balance function snippets and class implementations based on the operating systems random entropy (os.urandom). Decoupled tool metadata fields include tool_calls_json, tool_responses_json, and verification_payload, as independent JSON strings, which can be used for agent fine-tuning and reinforcement learning with programmatic rewards (RLVR) verification. The dataset size is approximately 100M to 1B samples, suitable for text generation, question answering, mathematical reasoning, code generation, and other scenarios.





