DICOM converted Slide Microscopy images for the CPTAC-OV collection
收藏资源简介:
This dataset corresponds to a collection of images and/or image-derived data available from National Cancer Institute Imaging Data Commons (IDC) [1]. This dataset was converted into DICOM representation and ingested by the IDC team. You can explore and visualize the corresponding images using IDC Portal here: CPTAC-OV. You can use the manifests included in this Zenodo record to download the content of the collection following the Download instructions below. Collection description This collection contains subjects from the National Cancer Institute’s Clinical Proteomic Tumor Analysis Consortium CPTAC Ovarian Serous Cystadenocarcinoma cohort. CPTAC is a national effort to accelerate the understanding of the molecular basis of cancer through the application of large-scale proteome and genome analysis, or proteogenomics. Please see the CPTAC-OV wiki page to learn more about the images and to obtain any supporting metadata for this collection. Files included A manifest file's name indicates the IDC data release in which a version of collection data was first introduced. For example, collection_id-idc_v8-aws.s5cmd corresponds to the contents of the collection_id collection introduced in IDC data release v8. If there is a subsequent version of this Zenodo page, it will indicate when a subsequent version of the corresponding collection was introduced. cptac_ov-idc_v10-aws.s5cmd: manifest of files available for download from public IDC Amazon Web Services buckets cptac_ov-idc_v10-gcs.s5cmd: manifest of files available for download from public IDC Google Cloud Storage buckets cptac_ov-idc_v10-dcf.dcf: Gen3 manifest (for details see https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids) Note that manifest files that end in -aws.s5cmd reference files stored in Amazon Web Services (AWS) buckets, while -gcs.s5cmd reference files in Google Cloud Storage. The actual files are identical and are mirrored between AWS and GCP. Download instructions Each of the manifests include instructions in the header on how to download the included files. To download the files using .s5cmd manifests: install idc-index package: pip install --upgrade idc-index download the files referenced by manifests included in this dataset by passing the .s5cmd manifest file: idc download manifest.s5cmd. To download the files using .dcf manifest, see manifest header. Acknowledgments Imaging Data Commons team has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under Task Order No. HHSN26110071 under Contract No. HHSN261201500003l. References [1] Fedorov, A., Longabaugh, W. J. R., Pot, D., Clunie, D. A., Pieper, S. D., Gibbs, D. L., Bridge, C., Herrmann, M. D., Homeyer, A., Lewis, R., Aerts, H. J. W., Krishnaswamy, D., Thiriveedhi, V. K., Ciausu, C., Schacherer, D. P., Bontempi, D., Pihl, T., Wagner, U., Farahani, K., Kim, E. & Kikinis, R. National Cancer Institute Imaging Data Commons: Toward Transparency, Reproducibility, and Scalability in Imaging Artificial Intelligence. RadioGraphics (2023). https://doi.org/10.1148/rg.230180
本数据集源自美国国家癌症研究所影像数据公共库(National Cancer Institute Imaging Data Commons,IDC)[1]收录的影像及影像衍生数据集合。本数据集已转换为DICOM格式,并由IDC团队完成数据摄入。您可通过IDC门户访问CPTAC-OV项目,浏览并可视化对应影像。您可遵循下方下载说明,使用本Zenodo记录中包含的清单文件下载该数据集集合的内容。 数据集集合说明 本集合包含美国国家癌症研究所临床蛋白质组肿瘤分析联盟(Clinical Proteomic Tumor Analysis Consortium,CPTAC)卵巢浆液性囊腺癌队列的受试对象数据。CPTAC是一项国家级项目,旨在通过大规模蛋白质组、基因组分析(即蛋白质基因组学)加速对癌症分子机制的认知。 如需了解更多影像相关信息并获取本数据集集合的配套元数据,请访问CPTAC-OV维基页面。 包含文件 清单文件的文件名可标识该集合数据版本首次发布所属的IDC数据发布版本。例如,collection_id-idc_v8-aws.s5cmd对应于IDC数据发布v8中首次推出的collection_id集合的内容。若本Zenodo页面有更新版本,将标注对应集合的后续版本发布时间。 cptac_ov-idc_v10-aws.s5cmd: 可从IDC公开Amazon Web Services(AWS)存储桶下载的文件清单 cptac_ov-idc_v10-gcs.s5cmd: 可从IDC公开Google Cloud Storage存储桶下载的文件清单 cptac_ov-idc_v10-dcf.dcf: Gen3格式清单(详情请访问https://learn.canceridc.dev/data/organization-of-data/guids-and-uuids) 请注意,后缀为-aws.s5cmd的清单文件指向存储于AWS存储桶的文件,而-gcs.s5cmd后缀的清单文件指向Google Cloud Storage中的文件。实际文件内容完全一致,且在AWS与GCP平台间实现了镜像同步。 下载说明 所有清单文件的头部均包含对应包含文件的下载指南。 使用.s5cmd格式清单下载文件: 1. 安装idc-index工具包:执行命令 pip install --upgrade idc-index 2. 通过传入.s5cmd格式的清单文件,下载本数据集中清单所指向的文件:idc download manifest.s5cmd 若使用.dcf格式清单下载文件,请参阅清单头部的说明。 致谢 影像数据公共库团队的研究经费全部或部分来自美国国家癌症研究所与美国国立卫生研究院的联邦拨款,项目编号为合同HHSN261201500003l下的任务订单HHSN26110071。 参考文献 [1] Fedorov A, Longabaugh W J R, Pot D, Clunie D A, Pieper S D, Gibbs D L, Bridge C, Herrmann M D, Homeyer A, Lewis R, Aerts H J W, Krishnaswamy D, Thiriveedhi V K, Ciausu C, Schacherer D P, Bontempi D, Pihl T, Wagner U, Farahani K, Kim E, Kikinis R. 美国国家癌症研究所影像数据公共库:实现影像人工智能研究的透明化、可复现性与可扩展性[J]. RadioGraphics, 2023. https://doi.org/10.1148/rg.230180



