materials-paper-corpus
收藏资源简介:
该数据集是一个从科学论文PDF中提取并构建的多模态语料库,旨在支持科学文献的信息提取、多模态内容理解、合成方法候选筛选以及自动化验证等任务。数据集包含多个迭代版本(v1, v3, v4, v5)和一个验证专用配置(corpus_v4_validation)。每个数据样本代表一篇科学论文,核心内容包括:论文元数据(如标题、摘要、作者、学科分类、出版日期、来源、DOI、PDF链接)、全文文本(从主论文和补充信息中提取的文本)、多模态内容(如图像路径和二进制数据,以及v3及以后版本中的图表元数据,包括图表ID、页码、标题、图注、边界框坐标、附近文本等)、处理状态与质量控制字段(记录提取成功状态、错误信息、PDF提取器、筛选状态)、评分与筛选字段(如合成评分、负面评分、图表正面评分、多模态优先级评分及筛选标志)、多模态增强信息(部分版本中的候选标志、摘要和状态),以及验证数据(专门用于验证提取结果,包含材料、结构、过程、设备、条件、语义准确性、格式合规性等多维评分、总体分数、验证通过标志和推理记录)。数据集按处理阶段分为多个分片,如元数据、增强后数据、合成候选、多模态增强、提取后数据和已验证提取结果,不同配置的数据规模从10个样本到3-10个样本不等,主要用于演示和测试目的。
This dataset is a multimodal corpus extracted and constructed from scientific paper PDFs, designed to support tasks such as information extraction from scientific literature, multimodal content understanding, synthesis method candidate screening, and automated validation. It includes multiple iterative versions (v1, v3, v4, v5) and a validation-specific configuration (corpus_v4_validation). Each data sample represents a scientific paper, with core features including: paper metadata (e.g., title, abstract, author list, subject classification, publication date, source, DOI, PDF link), full text (extracted from the main paper PDF as text_paper and from supporting information as text_si), multimodal content (such as images containing paths and binary byte data for embedded images, and figures in v3 and later versions with richer chart metadata like figure ID, page number, index, image path, caption, legend label, bounding box coordinates on the page, nearby text, extraction source, image dimensions), processing status and quality control fields (recording success status of text and figure extraction, error messages, PDF extractor used, filter passage status), scoring and filtering fields (including various scores for evaluation and filtering, such as synthesis score, negative score, figure positive score, multimodal priority score, and a passed_filter flag), multimodal enhancement information (in some versions, with candidate flags, summaries, enhancement status, and error messages), and validation data (specifically for validating extraction results, containing quantitative scores (0-1) across multiple dimensions like material extraction, structural integrity, process steps, equipment extraction, condition extraction, semantic accuracy, format compliance, along with overall scores, validation passage status, and reasoning process records). The dataset is divided into multiple splits based on processing stages, such as metadata, enriched, synthesis_candidates, multimodal_enriched, extracted, and validated_extractions. Different configurations vary in scale, ranging from 10 samples in the default configuration to 3-10 samples in others, primarily intended for demonstration and testing purposes.
数据集:materials-paper-corpus
该数据集是一个关于材料科学论文的多版本语料库,包含从论文中提取的文本、图像、图表及结构化元数据。
配置(Configurations)与规模
数据集包含6个配置,每个配置代表不同的处理阶段或子集:
- corpus_v1
- 特征:基本论文元数据、文本内容(
text_paper、text_si)、图像(images)、过滤标志及提取状态。 - 数据划分
- metadata: 10 条样本
- enriched: 10 条样本
- synthesis_candidates: 7 条样本
- 数据集大小:790,151 字节
- 特征:基本论文元数据、文本内容(
- corpus_v3
- 特征:在 v1 基础上新增提取的图表(
figures)详细信息(边界框、标题、附近文本等)、图表计数、处理标志及多模态富化相关字段。 - 数据划分
- metadata: 5 条样本
- enriched: 5 条样本
- synthesis_candidates: 4 条样本
- multimodal_enriched: 4 条样本
- 数据集大小:709,386 字节
- 特征:在 v1 基础上新增提取的图表(
- corpus_v4
- 特征:结构同 v3,但包含更新的处理状态和优先级评分。
- 数据划分
- metadata: 3 条样本
- enriched: 3 条样本
- synthesis_candidates: 3 条样本
- multimodal_enriched: 3 条样本
- extracted: 3 条样本
- 数据集大小:563,859 字节
- corpus_v4_validation
- 特征:专注于验证(
validation_status),包含从论文中提取的特定材料(material)及本体(ontology_extraction_json),并带有多种质量评分(如结构完整性、语义准确性)。 - 数据划分
- validated_extractions: 8 条样本
- 数据集大小:76,523 字节
- 特征:专注于验证(
- corpus_v5
- 特征:结构同 v4。
- 数据划分
- metadata: 4 条样本
- enriched: 4 条样本
- synthesis_candidates: 4 条样本
- multimodal_enriched: 4 条样本
- extracted: 4 条样本
- 数据集大小:761,585 字节
- default
- 特征:仅包含基本元数据、文本内容、图像及文本提取标志,无图表、多模态富化等高级特征。
- 数据划分
- metadata_test: 10 条样本
- 数据集大小:17,380 字节
总体数据规模
- 总下载大小:各配置下载大小之和为 2,227,534 字节。
- 总数据集大小:各配置数据集大小之和为 2,918,884 字节。




