Nemotron-VLM-Dataset-V2
收藏资源简介:
Nemotron-VLM-Dataset-V2是由英伟达构建的大规模多模态训练数据集,专为文档解析与OCR模型优化设计。该数据集融合合成数据、公共资源及人工标注样本,涵盖多语言文本、表格和密集OCR内容,数据量达数百万级,来源包括arXiv、Common Crawl及专业表格数据集。其创建通过创新管道实现,如NVpdftex工具链精准提取字符级边界框与语义标签,并辅以机器翻译和样式增强。该数据集主要应用于提升轻量级OCR模型在复杂文档理解、布局分析和多格式输出等领域的性能,旨在解决现代检索系统与语言模型对结构化文档信息的高精度提取需求。
Nemotron-VLM-Dataset-V2 is a large-scale multimodal training dataset developed by NVIDIA, specifically optimized for document parsing and OCR models. This dataset integrates synthetic data, publicly available resources, and manually annotated samples, covering multilingual text, tables, and dense OCR content, with a total of millions of samples. Its data sources include arXiv, Common Crawl, and professional table datasets. The construction of this dataset employs an innovative pipeline: the NVpdftex toolchain accurately extracts character-level bounding boxes and semantic labels, and the pipeline is further supplemented by machine translation and style augmentation. It is primarily designed to improve the performance of lightweight OCR models in complex document understanding, layout analysis, multi-format output and other related domains, aiming to address the demand for high-precision extraction of structured document information by modern retrieval systems and language models.

- 1NVIDIA Nemotron Parse 1.1英伟达 · 2025年



