PathOrchestra
收藏资源简介:
PathOrchestra数据集是由上海人工智能实验室等多个机构共同创建的,包含30万张来自20种不同组织和器官的病理切片的图像数据集。这些图像数据来源于三个中心的内部收藏和公共数据集。数据集被用于训练PathOrchestra模型,该模型通过自监督学习在无标签数据上学习高质量的特征表示,进而在有限的标注数据和参数下,在下游任务中表现出色。该数据集涵盖了数字切片预处理、全癌分类、病变识别、多癌亚型分类、生物标志物评估、基因表达预测和结构化报告生成等多种临床任务,旨在促进计算病理学领域的发展。
PathOrchestra dataset was jointly created by Shanghai AI Laboratory and multiple other institutions. It is a dataset containing 300,000 pathological slide images from 20 different tissues and organs. These image data are sourced from the internal collections of three centers and public datasets. The dataset is used to train the PathOrchestra model, which learns high-quality feature representations on unlabeled data via self-supervised learning, achieving excellent performance in downstream tasks with limited labeled data and model parameters. This dataset covers various clinical tasks including digital slide preprocessing, pan-cancer classification, lesion recognition, multi-cancer subtype classification, biomarker assessment, gene expression prediction and structured report generation, aiming to promote the development of the field of computational pathology.

- 1PathOrchestra: A Comprehensive Foundation Model for Computational Pathology with Over 100 Diverse Clinical-Grade Tasks上海人工智能实验室 · 2025年



