遇见数据集

BigMixSolDB: Extraction of a solubility database in solvent mixtures with an uncertainty-quantified large language model-based pipeline

收藏
Zenodo2026-04-02 更新2026-05-26 收录
官方服务:

资源简介:

Supplementary material for the publication "BigMixSolDB: Extraction of a solubility database in solvent mixtures with an uncertainty-quantified large language model-based pipeline". This Zenodo record contains the supplementary materials for BigMixSolDB, a workflow for extracting experimental solubility data in solvent mixtures from scientific PDFs. The archive includes a DOI list, prompt templates, request JSONL files, Docling- and VLM-based extraction outputs, standardized CSV datasets, and a source-code snapshot to support reproducibility of the extraction and validation pipeline. Files: dois.csv: List of publication DOIs included in the BigMixSolDB corpus. prompt_extraction.txt: Prompt template used to extract structured solubility records from parsed document text. prompt_training.txt: Prompt template used for fine-tuning. bigmixsoldb_docling.csv: Standardized BigMixSolDB dataset produced from the Docling-based PDF parsing pipeline. bigmixsoldb_vlm.csv: Standardized BigMixSolDB dataset produced from the vision-language-model-based PDF parsing pipeline. requests_docling.jsonl: Batch request log in JSONL format for extraction runs based on Docling-parsed inputs (Gemini 3 Pro models). requests_vlm.jsonl: Batch request log in JSONL format for extraction runs based on VLM-parsed inputs. extracted_vlm.zip: Archive of per-paper structured extraction outputs generated from the VLM pipeline. extracted_docling.zip: Archive of per-paper structured extraction outputs generated from the Docling pipeline. BigMixSolDB-main.zip: Snapshot of the BigMixSolDB source code and workflow used to generate, standardize, and evaluate the supplementary datasets. Also found on GitHub.

提供机构:
Zenodo
创建时间:
2026-04-02
二维码
社区交流群
二维码
科研交流群
商业服务