遇见数据集

BigMixSolDB: Extraction of a solubility database in solvent mixtures with an uncertainty-quantified large language model-based pipeline

收藏
Zenodo2026-05-27 更新2026-05-26 收录
官方服务:

资源简介:

Supplementary material for the publication "BigMixSolDB: Extraction of a solubility database in solvent mixtures with an uncertainty-quantified large language model-based pipeline". This Zenodo record contains the supplementary materials for BigMixSolDB, a workflow for extracting experimental solubility data in solvent mixtures from scientific PDFs. The archive includes a DOI list, prompt templates, request JSONL files, Docling- and VLM-based extraction outputs, standardized CSV datasets, and a source-code snapshot to support reproducibility of the extraction and validation pipeline. Files BigMixSolDB_code.zip: Snapshot of the BigMixSolDB source code and workflow used to generate, standardize, and evaluate the supplementary datasets. Also found on GitHub. bigmixsoldb_docling.csv: Standardized BigMixSolDB dataset produced from the Docling-based PDF parsing pipeline. Represents BigMixSolDB. bigmixsoldb_docling_unfiltered.csv: Unfiltered BigMixSolDB dataset produced from the Docling-based PDF parsing pipeline. bigmixsoldb_vlm.csv: Standardized BigMixSolDB dataset produced from the vision-language-model-based PDF parsing pipeline. bigmixsoldb_vlm_unfiltered.csv: Unfiltered BigMixSolDB dataset produced from the vision-language-model-based PDF parsing pipeline. doi_counts.csv: List of publication DOIs present in BigMixSolDB, BigSolDB 2.0, and MixtureSolDB, describing the number of entries present in each dataset. dois.csv: List of publication DOIs included retrieved during the creation the BigMixSolDB corpus. This represents all the sourced DOIs. excluded_solute_names.txt: List of solutes whose entries were excluded following a manual curation step. excluded_solvent_names.txt: List of solvents whose entries were excluded following a manual curation step. extracted_docling.zip: Archive of per-paper structured extraction outputs generated from the Docling pipeline. extracted_vlm.zip: Archive of per-paper structured extraction outputs generated from the VLM pipeline. llm_paper_screening_output.json: Screening results, employing Gemini 3.1 Pro Preview, of the DOIs containing ternary mixtures. Entries where the human reviewer disagreed are marked with the "human_decision" key. name_to_smiles_docling.json: List used to map molecule names to SMILES in the Docling-based dataset. name_to_smiles_vlm.json: List used to map molecule names to SMILES in the VLM-based dataset. prompt_extraction.txt: Prompt template used to extract structured solubility records from parsed document text. prompt_training.txt: Prompt template used for fine-tuning. removed_dois.txt: List of DOIs that were removed after the screening of literature containing ternary mixtures. requests_docling.jsonl: Batch request log in JSONL format for extraction runs based on Docling-parsed inputs (Gemini 3 Pro models). requests_vlm.jsonl: Batch request log in JSONL format for extraction runs based on VLM-parsed inputs. Changelog This version updates the Zenodo repository to match the final post-processed and filtered BigMixSolDB release described in the revised manuscript. Replaced the preliminary dataset files with the final filtered Docling-derived BigMixSolDB release containing 299,211 retained solubility entries from 2,301 source DOIs. Corrected the `review_required` flag, which had previously been incorrectly set to `False` for all rows in the deposited dataset. Standardized retained temperature, pressure, solubility, and solvent-composition units. Temperatures are reported in Kelvin, reported pressures in Pascal, solubilities and solvent compositions in fraction-based units. Manually curated solute and solvent name-to-SMILES mappings, corrected erroneous molecular identifiers, resolved ambiguous assignments, and removed rows with missing or invalid required SMILES. The raw extracted dataset can be found in `bigmixsoldb_docling_unfiltered.csv`. A list of molecules that were omitted in our dataset can be found in `excluded_solute_names.txt` and `excluded_solvents.txt`. Post-processing steps and filtering the molecules found in the aforementioned TXT files gives the final dataset present in `bigmixsoldb_docling.csv`. Applied final scope filtering to retain molecular solutes in single- and multi-component molecular solvent systems, removing polymeric or ill-defined systems, ionic liquids, deep eutectic solvents, inorganic molten-salt systems, and solvent systems with more than three components. Added supporting files documenting per-DOI coverage comparisons against BigSolDB 2.0 and MixtureSolDB, list of excluded molecules from the final dataset, and ternary-mixture source screening. Updated summary statistics for the final release: 1,837 unique solutes, 450 unique solvents, 1,020 binary solvent systems, and 141 ternary solvent systems.

本数据集为发表论文"BigMixSolDB:基于带不确定性量化的大语言模型(Large Language Model,LLM)管线构建溶剂混合物溶解度数据库"的补充材料。 本Zenodo存档包含BigMixSolDB相关补充材料,该工作流用于从学术PDF文档中提取溶剂混合物的实验溶解度数据。本存档包涵盖DOI列表、提示词模板、请求用JSONL格式文件、基于Docling与视觉语言模型(Vision-Language Model,VLM)的提取结果、标准化CSV格式数据集,以及用于支撑提取与验证管线可复现性的源代码快照。 文件列表: dois.csv:收录于BigMixSolDB语料库的学术文献DOI列表。 prompt_extraction.txt:用于从解析后的文档文本中提取结构化溶解度记录的提示词模板。 prompt_training.txt:用于模型微调的提示词模板。 bigmixsoldb_docling.csv:基于Docling的PDF解析管线生成的标准化BigMixSolDB数据集。 bigmixsoldb_vlm.csv:基于VLM的PDF解析管线生成的标准化BigMixSolDB数据集。 requests_docling.jsonl:基于Docling解析结果(使用Gemini 3 Pro模型)的提取任务的JSONL格式批量请求日志。 requests_vlm.jsonl:基于VLM解析结果的提取任务的JSONL格式批量请求日志。 extracted_vlm.zip:由VLM管线生成的单篇论文结构化提取结果存档包。 extracted_docling.zip:由Docling管线生成的单篇论文结构化提取结果存档包。 BigMixSolDB-main.zip:用于生成、标准化与评估本次补充数据集的BigMixSolDB源代码及工作流快照,该代码亦托管于GitHub平台。

提供机构:
Zenodo
创建时间:
2026-04-02
二维码
社区交流群
二维码
科研交流群
商业服务