SmolDocling-256M-preview
收藏资源简介:
SmolDocling-256M-preview是IBM Research和HuggingFace共同构建的一个文档理解数据集,该数据集为公开源数据集,包含了图表、表格、公式和代码识别等任务所需的数据。数据集由来自CommonCrawl、Wikipedia和商业相关文档的PDF文档构成,通过一系列处理步骤增加了弱标注信息,以支持文档元素如方程、表格、代码和图表的视觉识别和结构化理解。该数据集旨在促进文档转换和结构识别任务的研究,解决文档理解中的格式多样性和布局复杂性挑战。
SmolDocling-256M-preview is a document understanding dataset jointly developed by IBM Research and HuggingFace. As an open-access public dataset, it includes data required for tasks such as chart recognition, table recognition, formula recognition and code recognition. The dataset is constructed from PDF documents sourced from CommonCrawl, Wikipedia and commercially relevant documents, with a series of processing steps applied to add weakly annotated information to support visual recognition and structured understanding of document elements including equations, tables, code and charts. This dataset aims to promote research on document conversion and structure recognition tasks, and address the challenges of format diversity and layout complexity in document understanding.

- 1SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversionIBM Research, HuggingFace · 2025年



