Supporting data for "JUMP Cell Painting chemical dataset: Morphological map of 116k compounds in human cells"
收藏资源简介:
Exact input snapshots used by the jpx and jump_production analysis pipelines supporting "JUMP Cell Painting chemical dataset: Morphological map of 116k compounds in human cells." File descriptions targetannotations_singleproteins_jumpcompounds_all_5.csv - Wide ChEMBL release 33 compound-target activity matrix for 17,892 JUMP compounds and 102 single-protein targets identified by UniProt accession. The file contains the JUMP compound ID and standardized SMILES, with 1 for active, 0 for inactive, and blank for an untested or unretained compound-target pair. It contains 24,029 active and 44,182 inactive pairs. The pipeline retains active links, combines each compound's UniProt accessions, and loads them into the chembl_protein_targets table. toxicity_pk_data.zip - Bundle of three standardized-SMILES datasets: DILIrank v2 drug-induced liver injury concern levels (1,146 compounds), DICTrank drug-induced cardiotoxicity concern levels (1,235 compounds), and PKSmart human pharmacokinetic measurements (1,283 compounds). The PK endpoints are steady-state volume of distribution in L/kg, clearance in mL/min/kg, fraction unbound in plasma, mean residence time in hours, and half-life in hours. The pipeline matches these records to JUMP compounds by the InChIKey connectivity layer and loads them into toxicity_pk_annotations. 2026-04-06_JUMP_chems_func_use.csv - JUMP-specific EPA functional-use annotations for 10,451 compounds mapped through DSSTox substance identifiers. Each row includes JUMP and chemical identifiers, curated CPDat functional-use categories, and sparse model-predicted probabilities from 0 to 1 for 35 industrial and consumer use categories following Phillips et al. (2017). Curated categories may be multi-label; predicted probabilities are model outputs rather than confirmed uses, and NA means missing rather than zero. The pipeline loads the cleaned values into the functional_use table. pd_export_01_2025_875_targets_standardized.xlsx - January 2025 manual export from the Probes & Drugs portal. The COMPOUNDS sheet contains 875 chemical probes with structures, identifiers, probe and drug-status flags, quality alerts, and physicochemical properties. The TARGETS sheet contains 6,641 probe-target relationships with gene, target, mechanism-of-action, activity, and protein-classification fields. The pipeline standardizes and matches the probes to JUMP compounds and loads compound-level and detailed target tables. pkis_chembl_cache.json - Frozen cache consumed by the kinase-probe processing script, containing 2,000 raw ChEMBL molecule API records with identifiers, structures, properties, synonyms, and status flags; 1,968 records contain molecular structures. The filename is historical and this file should be treated as the exact API cache used by the pipeline, not as a standalone authoritative PKIS membership table. mitotox_compounds.parquet - Derived snapshot of 1,383 compounds from the MitoTox API with canonical SMILES obtained from PubChem. Columns contain the source compound name and identifiers, a toxic or non-toxic label, and affected mitochondrial functions. A compound is labeled toxic when at least one source experimental record is positive; positive records are summarized across eight top-level mitochondrial function categories. The pipeline matches the compounds to JUMP and loads the labels into mitotox_annotations. molport_batch_search.zip - Snapshot of a manual MolPort commercial-availability search for the JUMP compound collection. It contains 115,796 submitted SMILES divided among 12 batch files and 12 quote spreadsheets with 72,363 matched search rows, including MolPort ID, match type, supplier, catalog number, package, and quoted price. The analysis uses this snapshot for exact-SMILES commercial-availability filtering, including a USD 200 price cutoff; it is not a live inventory or pricing feed. jump_to_chembl.csv.gz - Checksum-pinned ranked SmallWorld exact-structure search output mapping 28,782 JUMP compound identifiers to 36,449 ChEMBL molecule identifiers in 36,488 rows. It was generated against chembl_31.anon and preserves 7,706 non-rank-zero alternatives required to reproduce the published ChEMBL 33 target-annotation matrix. The generator verifies its SHA-256 digest and derives the compound roster and ChEMBL-to-UniProt map from data/external/compound.csv.gz and official ChEMBL 33 SQLite, respectively. Canonical external input pkis2.xlsx is not duplicated in this deposit because the required file is byte-identical to the Drewry et al. (2017) PLOS ONE S4 supplement at doi:10.1371/journal.pone.0181585.s004. It contains 646 PKIS2 compounds, compound metadata, and percentage inhibition measurements at 1 micromolar across 406 kinase columns. Version 4 adds the three jump_production inputs that lacked canonical archival URLs and expands the descriptions for all retained files. Version 5 adds the sole irreducible historical compatibility artifact, the ranked SmallWorld mapping. The two other former intermediate snapshots are now derived from canonical sources by the repository generator.



