HoneyBee
收藏资源简介:
HoneyBee是一个用于创建多模态肿瘤学数据集的可扩展模块化框架,由莫菲特癌症中心开发。该数据集整合了临床记录、影像数据和患者结果等多种数据模态,通过基础模型生成代表性嵌入。数据集大小庞大,包含来自TCGA项目的11,428名患者的数据,涵盖33种癌症类型。创建过程中,利用了先进的数据预处理技术和基于变压器的架构来生成嵌入,捕捉原始医疗数据中的基本特征和关系。HoneyBee旨在通过提供高质量、机器学习就绪的数据集,加速肿瘤学研究,解决医疗数据复杂性和异质性的挑战,并可扩展到其他医疗领域。
HoneyBee is a scalable, modular framework for creating multimodal oncology datasets, developed by the Moffitt Cancer Center. This dataset integrates multiple data modalities including clinical records, imaging data, and patient outcomes, and generates representative embeddings via foundation models. Boasting a large scale, the dataset contains data from 11,428 patients in the TCGA project, covering 33 cancer types. During its development, advanced data preprocessing techniques and Transformer-based architectures are employed to generate embeddings that capture the fundamental features and relationships within raw medical data. HoneyBee aims to accelerate oncology research by providing high-quality, machine learning-ready datasets, address the challenges of complexity and heterogeneity in medical data, and be scalable to other healthcare domains.




