DaTikZ-V4
收藏资源简介:
DaTikZ-V4 是一个用于训练 TikZ/LaTeX 图表生成模型(如 TikZilla-3B、TikZilla-8B 等)的数据集,旨在从自然语言描述生成 TikZ/LaTeX 科学图表。数据集包含 10 万到 100 万条样本,数据来源于 ArXiv、GitHub 和 TeXStackExchange 的 TikZ 代码,并使用 Qwen2.5-VL-7B-Instruct 生成对应的科学图表描述。每个样本包含以下字段:唯一标识符 (file_id)、原始标题 (caption)、视觉语言模型生成的详细描述 (vlm_description)、完整的 LaTeX/TikZ 源代码 (tikz_code)、数据来源 (source) 以及渲染后的图表图像 (png_image)。该数据集适用于文本生成任务,特别是科学图表生成、代码生成和 LaTeX 相关应用。
DaTikZ-V4 is a dataset designed for training TikZ/LaTeX diagram generation models (e.g., TikZilla-3B, TikZilla-8B) to generate scientific TikZ/LaTeX diagrams from natural language descriptions. The dataset contains 100,000 to 1,000,000 samples, with source data collected from TikZ codes on ArXiv, GitHub, and TeXStackExchange, and the corresponding scientific diagram descriptions generated using Qwen2.5-VL-7B-Instruct. Each sample includes the following fields: unique identifier (file_id), original caption, detailed description generated by the vision-language model (vlm_description), complete LaTeX/TikZ source code (tikz_code), data source (source), and rendered diagram image (png_image). This dataset is suitable for text generation tasks, especially scientific diagram generation, code generation, and LaTeX-related applications.




