遇见数据集

LockOnN/chart-think-with-images-sft

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

一个合成的SFT数据集,用于训练多模态大型语言模型(LLMs)在图表图像上进行工具增强推理。模型学习在关于图表的思维链推理中使用`crop`(图像区域缩放)和`code_interpreter`(Python代码执行)工具。数据集包含643个示例,来源于ChartQA(76.4%)和CharXiv(23.6%)两个数据集,格式为ChatML消息(系统+用户+助手)。数据集特别关注模型在推理过程中使用工具的能力,包括何时裁剪、何时计算、何时组合工具以及错误恢复。数据集的类别分布包括`spatial_detail`(33.6%)、`visual_lookup`(23.5%)、`calculation`(21.5%)、`multi_tool`(11.7%)和`error_correction`(9.8%)。每个示例包含ChatML格式的消息,其中包含工具增强的推理过程。

A synthetic Supervised Fine-Tuning (SFT) dataset designed for training multimodal Large Language Models (LLMs) to perform tool-augmented reasoning over chart images. The model learns to utilize two tools, `crop` (image region zooming) and `code_interpreter` (Python code execution), during chain-of-thought reasoning about charts. This dataset consists of 643 examples, sourced from two existing datasets: ChartQA (76.4%) and CharXiv (23.6%), and is formatted as ChatML messages (system + user + assistant). This dataset specifically focuses on the model's ability to use tools during reasoning, including when to crop, when to perform calculations, when to combine multiple tools, and error recovery. The category distribution of the dataset includes `spatial_detail` (33.6%), `visual_lookup` (23.5%), `calculation` (21.5%), `multi_tool` (11.7%), and `error_correction` (9.8%). Each example contains ChatML-formatted messages that encapsulate tool-augmented reasoning processes.

提供机构:
LockOnN
二维码
社区交流群
二维码
科研交流群
商业服务