2026_CanResComm
收藏资源简介:
Manuscript Metadata Manuscript title: Drug screening of sarcoma cells: finding shared sensitivities Corresponding author: Dr. Beird Authors: Hannah Beird, Carl Ho, Roberto Cardenas-Zuniga, Asmaa Ahmed, Clement Agyemang, Danh Truong, Clifford Stephan, Yongsung Park, Reid Powell, Kevin Murgas, Joseph Tym, Sharon Landers, Stephanie Schmidt, Noha Osman, Dejka Araujo, Anthony Conley, Neeta Somaiah, Bissan Al-Lazikani, Joseph Ludwig, and P. Andrew Futreal Package Purpose This Zenodo package preserves the screening archive that supports the manuscript above. It contains the raw plate-level count data, the platemap definitions that describe compound placement and concentrations, the QC reports that benchmark every run, the leveled analytic outputs, and the campaign-level summaries referenced in the paper. Package Layout 01_metadata/manifest_canonical.tsv: canonical manifest with linkages from plate barcode to platemap, run date, canonical cell-line name, and Cellosaurus accession when available. 01_metadata/cell_line_catalog.tsv: canonical cell-line catalog with supplied identifiers and observed plate counts. 01_metadata/legacy_alias_crosswalk.tsv: mapping from legacy archive cell-line aliases to canonical labels for provenance. 01_metadata/compound_catalog_with_public_ids.tsv: aggregated platemap catalog enriched with PubChem CID, InChIKey, SMILES, and formula where a single automated resolution was possible. 01_metadata/platemap_well_annotations.tsv.gz: compressed, well-level platemap annotations (Library, Row, Col, Func, cmpd1, Conc1) parsed directly from each Excel workbook. 02_raw_data/raw_counts.zip: all raw growth/count CSVs reorganized under raw_counts/<canonical_cell_line>/<run>/. 02_raw_data/plate_maps.zip: the original platemap Excel files referenced by the manifest. 03_processed_data/quality_control.zip: per-cell-line/run QC outputs with growth charts, Z-prime, CV, and reference-agent plots. 03_processed_data/leveled_analysis.zip: leveled data organized by canonical cell line, including Lvl_1, Lvl_2, and Sing_DRC folders. 03_processed_data/campaign_summary.zip: campaign-level summary tables and heatmaps for cross-run review. 04_analysis_scripts/: supplementary R script and input tables used to generate the manuscript-level filtering and UpSet overlap analyses. SHA256SUMS.tsv: checksums for every file in this lean package. Archive Summary Plates in manifest: 1,050 Platemap workbooks: 848 Raw CSV files: 1,050 Distinct run dates: 42 Distinct archive cell-line labels: 30 Distinct canonical cell-line names after normalization: 25 Unique compound names evaluated for public identifiers: 1,403 Compound names resolved to a single PubChem CID: 1,364 Compound names with multiple CID matches left unresolved: 29 Compound names unresolved or not applicable: 10 Major platemap libraries observed: Selleck2, Selleck1, Selleck3 Data Model and Provenance 1. Manifest The manifest links every screened plate to its canonical cell line, platemap, run date, and supplied identifier information. Key fields include: plate_barcode platemap_id screen_date canonical_cell_line identifier_supplied_in_chat cellosaurus_url set run 2. Raw Data Raw CSVs contain columns Row, Col, and count corresponding to per-well readouts from the cell-count-based assay. Files are organized by canonical cell line to simplify navigation while retaining the original file names. 3. Plate Maps and Compound Identifiers Each platemap workbook exposes Library, Row, Col, Func, cmpd1 (compound name), and Conc1 (compound concentration). These annotations feed into platemap_well_annotations.tsv.gz, which is then enriched by PubChem metadata in compound_catalog_with_public_ids.tsv. Only compounds with a single unambiguous CID (including conservative normalized query fallbacks) were filled automatically; ambiguous or unresolvable names remain empty to avoid speculation. 4. Processed Results This package preserves the processed outputs that directly support the paper: quality_control.zip: run- and cell-line-specific QC visualizations and tables. leveled_analysis.zip: leveled data layers and fitted curves grouped by canonical cell line. campaign_summary.zip: campaign-wide tables and heatmaps for cross-run comparison. 5. Additional Analysis Scripts The additional analysis scripts and data added after the initial curation are included intact: Gulf-Coast-filter-forZenodo.R GCC_SARC_ALl_AUC_FA-2016_inputforfilter-250815.txt GCC-SARC-top10perc-quantile-Upset-input-260316.txt GULF-COAST-HISTOLOGY.txt These resources were used to filter the AUC matrix, generate heatmaps with histology annotations, and create the UpSet-style overlap summaries referenced in the manuscript. The R script retains the original ~/Documents/... paths and will require updating if re-run in another environment. Assay QC The QC archive is the authoritative record: each cell line/run includes images and tables for growth, Z-prime, control CV, and reference-agent dose responses, plus a tabular workbook such as 6) <cellline>_<run>_Tabular Data.xlsx. Reuse Guidance Use 01_metadata/manifest_canonical.tsv to map each plate barcoded CSV to its canonical cell line, run, and Cellosaurus accession. Use 01_metadata/platemap_well_annotations.tsv.gz to join well positions to compounds and concentrations. Use 01_metadata/compound_catalog_with_public_ids.tsv when PubChem identifiers (CID, InChIKey, SMILES) are required. Join raw counts to platemap entries by well position before performing analyses. Consult 03_processed_data/quality_control.zip and 03_processed_data/campaign_summary.zip for QC verification. Use the leveled outputs in 03_processed_data/leveled_analysis.zip when reproducing fitted single-agent summaries. Explore 04_analysis_scripts/ for the supplementary filtering and overlap analyses; update absolute paths before re-running the script.



