遇见数据集

BigMixSolDB: Extraction of a solubility database in solvent mixtures with an uncertainty-quantified large language model-based pipeline

收藏
Zenodo2026-05-27 更新2026-05-29 收录
官方服务:

资源简介:

Supplementary material for the publication "BigMixSolDB: Extraction of a solubility database in solvent mixtures with an uncertainty-quantified large language model-based pipeline". This Zenodo record contains the supplementary materials for BigMixSolDB, a workflow for extracting experimental solubility data in solvent mixtures from scientific PDFs. The archive includes a DOI list, prompt templates, request JSONL files, Docling- and VLM-based extraction outputs, standardized CSV datasets, and a source-code snapshot to support reproducibility of the extraction and validation pipeline. Files BigMixSolDB_code.zip: Snapshot of the BigMixSolDB source code and workflow used to generate, standardize, and evaluate the supplementary datasets. Also found on GitHub. bigmixsoldb_docling.csv: Standardized BigMixSolDB dataset produced from the Docling-based PDF parsing pipeline. Represents BigMixSolDB. bigmixsoldb_docling_unfiltered.csv: Unfiltered BigMixSolDB dataset produced from the Docling-based PDF parsing pipeline. bigmixsoldb_vlm.csv: Standardized BigMixSolDB dataset produced from the vision-language-model-based PDF parsing pipeline. bigmixsoldb_vlm_unfiltered.csv: Unfiltered BigMixSolDB dataset produced from the vision-language-model-based PDF parsing pipeline. doi_counts.csv: List of publication DOIs present in BigMixSolDB, BigSolDB 2.0, and MixtureSolDB, describing the number of entries present in each dataset. dois.csv: List of publication DOIs included retrieved during the creation the BigMixSolDB corpus. This represents all the sourced DOIs. excluded_solute_names.txt: List of solutes whose entries were excluded following a manual curation step. excluded_solvent_names.txt: List of solvents whose entries were excluded following a manual curation step. extracted_docling.zip: Archive of per-paper structured extraction outputs generated from the Docling pipeline. extracted_vlm.zip: Archive of per-paper structured extraction outputs generated from the VLM pipeline. llm_paper_screening_output.json: Screening results, employing Gemini 3.1 Pro Preview, of the DOIs containing ternary mixtures. Entries where the human reviewer disagreed are marked with the "human_decision" key. name_to_smiles_docling.json: List used to map molecule names to SMILES in the Docling-based dataset. name_to_smiles_vlm.json: List used to map molecule names to SMILES in the VLM-based dataset. prompt_extraction.txt: Prompt template used to extract structured solubility records from parsed document text. prompt_training.txt: Prompt template used for fine-tuning. removed_dois.txt: List of DOIs that were removed after the screening of literature containing ternary mixtures. requests_docling.jsonl: Batch request log in JSONL format for extraction runs based on Docling-parsed inputs (Gemini 3 Pro models). requests_vlm.jsonl: Batch request log in JSONL format for extraction runs based on VLM-parsed inputs. Changelog This version updates the Zenodo repository to match the final post-processed and filtered BigMixSolDB release described in the revised manuscript. Replaced the preliminary dataset files with the final filtered Docling-derived BigMixSolDB release containing 299,211 retained solubility entries from 2,301 source DOIs. Corrected the `review_required` flag, which had previously been incorrectly set to `False` for all rows in the deposited dataset. Standardized retained temperature, pressure, solubility, and solvent-composition units. Temperatures are reported in Kelvin, reported pressures in Pascal, solubilities and solvent compositions in fraction-based units. Manually curated solute and solvent name-to-SMILES mappings, corrected erroneous molecular identifiers, resolved ambiguous assignments, and removed rows with missing or invalid required SMILES. The raw extracted dataset can be found in `bigmixsoldb_docling_unfiltered.csv`. A list of molecules that were omitted in our dataset can be found in `excluded_solute_names.txt` and `excluded_solvents.txt`. Post-processing steps and filtering the molecules found in the aforementioned TXT files gives the final dataset present in `bigmixsoldb_docling.csv`. Applied final scope filtering to retain molecular solutes in single- and multi-component molecular solvent systems, removing polymeric or ill-defined systems, ionic liquids, deep eutectic solvents, inorganic molten-salt systems, and solvent systems with more than three components. Added supporting files documenting per-DOI coverage comparisons against BigSolDB 2.0 and MixtureSolDB, list of excluded molecules from the final dataset, and ternary-mixture source screening. Updated summary statistics for the final release: 1,837 unique solutes, 450 unique solvents, 1,020 binary solvent systems, and 141 ternary solvent systems.

提供机构:
Zenodo
创建时间:
2026-05-27
二维码
社区交流群
二维码
科研交流群
商业服务