diagram_desc_datikz
收藏资源简介:
该数据集包含58,238个训练样本,每个样本由四个文本字段构成:file_id(文件标识符)、tikz_code(TikZ绘图代码)、reasoning_trace(推理轨迹文本)和generated_description(生成的描述文本)。数据以纯文本形式组织,适用于代码生成、多模态推理、文本到图形代码的转换等任务,尤其可能用于研究TikZ代码的自动生成与理解、多步推理过程建模或描述与代码的对齐学习。
This dataset contains 58,238 training samples, each consisting of four text fields: file_id (file identifier), tikz_code (TikZ drawing code), reasoning_trace (reasoning trace text), and generated_description (generated description text). The data is organized in plain text format and is suitable for tasks such as code generation, multimodal reasoning, and text-to-graphics code conversion. It is particularly useful for research on automatic generation and understanding of TikZ code, modeling of multi-step reasoning processes, or alignment learning between descriptions and code.




