thainamhoang/ViMed-PET-CT
收藏资源简介:
--- license: cc-by-4.0 task_categories: - image-to-text - text-generation - image-text-to-text size_categories: - 1K<n<10K --- # ViMed-PET-CT Forked and optimized compression of [dacthai2807/ViMed-PET](https://huggingface.co/datasets/dacthai2807/ViMed-PET), converting `.npy` and chunked zip files into `.npz` files. Better annotation and guideline. Data includes **2017, 2018, 2019, and 2023**. Each patient contains: - basic metadata: sex, height, weight - CT scan - PET scan - generated report A separate `medical_test_set/` folder is included. ## Year Coverage - 2017: August to December - 2018: all year except May and June - 2019: May, June, October, November, December - 2023: whole year ## Scan Shape - CT: `(313, 512, 512)` - PET: `(313, 256, 256)` ## Metadata `metadata.csv` is used for dataset preview. It contains: - sex - height - weight - year - direct path to PET - direct path to CT - direct path to reports ## Citation ``` @misc{nguyen2026visionlanguagefoundationmodelmedical, title={Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation}, author={Huu Tien Nguyen and Dac Thai Nguyen and The Minh Duc Nguyen and Trung Thanh Nguyen and Thao Nguyen Truong and Huy Hieu Pham and Johan Barthelemy and Minh Quan Tran and Thanh Tam Nguyen and Quoc Viet Hung Nguyen and Quynh Anh Chau and Hong Son Mai and Thanh Trung Nguyen and Phi Le Nguyen}, year={2026}, eprint={2509.24739}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2509.24739}, } ```
许可证:CC BY 4.0 任务类别: - 图像到文本(image-to-text) - 文本生成(text-generation) - 图像文本到文本(image-text-to-text) 数据规模:1000 < 样本数量 < 10000 # ViMed-PET-CT 本数据集系对[dacthai2807/ViMed-PET](https://huggingface.co/datasets/dacthai2807/ViMed-PET)的复刻与压缩优化版本,将原有的`.npy`与分块zip文件转换为`.npz`格式文件。 优化了标注规范与使用指南。 数据集涵盖2017、2018、2019及2023年的样本,每位患者的数据包含: - 基础元数据:性别、身高、体重 - CT扫描影像 - PET扫描影像 - 生成的医学报告 数据集额外包含独立的`medical_test_set/`测试集文件夹。 ## 数据覆盖时段 - 2017年:8月至12月 - 2018年:除5月、6月外的全年时段 - 2019年:5月、6月、10月、11月、12月 - 2023年:全年时段 ## 影像维度 - CT影像:`(313, 512, 512)` - PET影像:`(313, 256, 256)` ## 元数据 可通过`metadata.csv`文件快速预览数据集的整体信息,该文件包含以下字段: - 性别 - 身高 - 体重 - 采样年份 - PET影像直接存储路径 - CT影像直接存储路径 - 医学报告直接存储路径 ## 引用格式 bibtex @misc{nguyen2026visionlanguagefoundationmodelmedical, title={面向医疗数据的视觉语言基础模型:越南语PET/CT报告生成的多模态数据集与基准}, author={Huu Tien Nguyen and Dac Thai Nguyen and The Minh Duc Nguyen and Trung Thanh Nguyen and Thao Nguyen Truong and Huy Hieu Pham and Johan Barthelemy and Minh Quan Tran and Thanh Tam Nguyen and Quoc Viet Hung Nguyen and Quynh Anh Chau and Hong Son Mai and Thanh Trung Nguyen and Phi Le Nguyen}, year={2026}, eprint={2509.24739}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2509.24739}, }




