Lightcap/nocisnn-turkish-logic-traces
收藏资源简介:
NociSNN土耳其数学逻辑轨迹数据集记录了由土耳其语LLM模型在持续运行的NociSNN评估循环中生成的结构化数学推理轨迹。数据生成基于确定性土耳其语模板,涵盖算术、百分比、比例和单变量线性方程问题,真实答案由任务生成器计算。模型为每个问题生成简短的JSON解决步骤,验证器检查格式、与先前步骤的链接、重复、漂移、矛盾及最终答案正确性。数据集保存了接受和拒绝的尝试,以便研究惩罚行为。每条JSONL记录包含合成土耳其语问题、代码导出的预期答案、模型标识符、策略摘要、结构化步骤(或原始输出字符串)、验证标志(如接受、重复、漂移、矛盾、最终正确)以及NociSNN观察值和状态摘要。该数据集是自动生成的研究观察数据,非人工编写的数学解释或监督黄金推理,使用者应利用包含的正确性和拒绝元数据,避免将所有行视为正面训练目标。
The NociSNN Turkish Mathematical Logic Trace Dataset documents structured mathematical reasoning traces generated by Turkish LLMs during the continuously running NociSNN evaluation loop. Data generation is based on deterministic Turkish templates, covering arithmetic, percentage, ratio, and univariate linear equation problems, with ground-truth answers calculated by the task generator. For each problem, the model generates concise JSON-formatted solution steps, while a validator checks formatting compliance, links to prior steps, repetitions, drift, contradictions, and the correctness of the final answer. The dataset stores both accepted and rejected attempts to support research on penalty behaviors. Each JSONL entry contains a synthesized Turkish question, code-derived expected answer, model identifier, policy summary, structured steps (or raw output string), validation flags (e.g., "accepted", "repeated", "drifted", "contradictory", "finally correct"), as well as NociSNN observations and state summaries. This dataset consists of automatically generated research observation data, rather than manually written mathematical explanations or supervised gold-standard reasoning. Users should leverage the included correctness and rejection metadata, and avoid treating all entries as positive training targets.




