遇见数据集

IPATH Dataset: 45,609 Curated Image-Text Pairs for Histopathology Applications

收藏
Zenodo2026-01-27 更新2026-05-26 收录
官方服务:

资源简介:

Recent advancements in artificial intelligence (AI) have revealed important patterns in pathology images imperceptible to human observers that can improve diagnostic accuracy and decision support systems. However, progress has been limited due to the lack of publicly available medical images. To address this scarcity, we explore Instagram as a novel source of pathology images with expert annotations. We curated the IPATH dataset from Instagram, comprising 45,609 pathology image–text pairs rigorously filtered and curated for domain quality using classifiers, large language models, and manual filtering. To demonstrate the value of this dataset, we developed a multimodal AI model called IP-CLIP by fine-tuning a pretrained CLIP model using the IPATH dataset. We evaluated IP-CLIP on seven external histopathology datasets using zero shot classification and linear probing, where it consistently outperformed the original CLIP model. Furthermore, IP-CLIP matched or exceeded several recent state-of-the-art pathology vision–language models, despite being trained on a substantially smaller dataset. We also assessed image–text alignment on a 5k held-out IPATH subset using image–text retrieval, where IP-CLIP surpassed CLIP and other specialized models. These results demonstrate the effectiveness of the IPATH dataset and highlight the potential of leveraging social media data to develop AI models for medical image classification and enhance diagnostic accuracy. What's New in Version 2.0 This is the cleaned and preprocessed version of the IPATH dataset, ready for direct use in training vision-language models. Compared to v1.0 (raw images), v2.0 includes: ✅ 1. Circular Field-of-View (FOV) Correction Many microscope images contain circular viewing areas with black borders. V2.0 applies automated correction using: RANSAC-based circle fitting for robust detection (handles partial circles) Intelligent cropping to remove black borders Background filling with tissue-appropriate colors Works on both complete and partial circular FOVs that touch image edges Impact: Removes distracting artifacts and focuses model attention on tissue content. ✅ 2. Text and Annotation Removal Clinical pathology images often contain patient identifiers, scale bars, magnification markers, and other text overlays. V2.0 removes these using: OCR-based detection (PaddleOCR) for precise text localization Conservative fallback detection for text OCR might miss Inpainting to seamlessly fill removed text regions Privacy protection by automatically removing any visible text Impact: Protects patient privacy and prevents models from learning text artifacts instead of tissue patterns. ✅ 3. Standardized Format and Resolution All images in v2.0 are standardized to: Format: PNG (lossless compression, no JPEG artifacts) Resolution: 512×512 pixels Preprocessing: Shortest side resized to 512px, maintaining aspect ratio Quality: High-quality images suitable for model training Impact: Consistent input format for reproducible model training.

提供机构:
Zenodo
创建时间:
2026-01-27
二维码
社区交流群
二维码
科研交流群
商业服务