sionic-ai/mdpbench-doc-ocr-sft
收藏资源简介:
mdpbench-doc-ocr-sft数据集是一个用于Qwen 4B(Qwen3-VL)模型的监督微调学习数据,专注于韩语和日语文档的OCR(光学字符识别)和解析任务。数据集包含韩国地方政府新闻通讯(报纸形式)的图像和Markdown格式的真实文本,旨在支持文档图像到文本的转换,并保留表格和阅读顺序。数据来源包括韩国地方政府出版物,如Seongnam Vision Seongnam、Seocho-gu和Gongju-si,这些是公共著作物,允许非商业和研究用途。数据集确保无污染,与MDPBench测试集无交集,并提供了训练示例和加载方法。
The mdpbench-doc-ocr-sft dataset is a supervised fine-tuning (SFT) dataset for the Qwen 4B (Qwen3-VL) model, focusing on OCR (Optical Character Recognition) and parsing tasks for Korean and Japanese documents. The dataset contains images and Markdown-formatted ground truth texts of Korean local government newsletters (in newspaper format), aiming to support document image-to-text conversion while preserving tables and reading order. Its data sources include Korean local government publications such as Seongnam Vision Seongnam, Seocho-gu, and Gongju-si, which are public domain works permitting non-commercial and research use. The dataset is guaranteed to be uncontaminated, with no overlap with the MDPBench test set, and provides training examples and loading methods.




