遇见数据集

Tcruz-fpn dataset: slide level partitioned audrant images for Trypanosoma cruzi classification from the Morais et al. smartphone microscopy corpus

收藏
Zenodo2026-04-23 更新2026-05-26 收录
官方服务:

资源简介:

## Contents ``` dataset_split_fieldlevel/ ├── README.md # this file ├── manifest.txt # summary of the split: seed, field ranges per partition ├── train/ │ ├── positive_images/ # 2,528 PNGs │ └── negative_images/ # 2,528 PNGs ├── val/ │ ├── positive_images/ # 188 PNGs │ └── negative_images/ # 124 PNGs └── test/ ├── positive_images/ # 279 PNGs └── negative_images/ # 141 PNGs ``` All PNG files are 1,224 × 1,632 pixel lossless RGB. Total: 5,056 training images (flip-balanced + Gaussian-noise-augmented to a 1:1 positive/negative ratio), 312 validation images, 420 test images (validation and test are unaugmented). File naming convention: `field{NNNN}_quad{K}[.augmentation_tag].png`, where `NNNN` is the Morais field number (0001–0704) and `K` is the quadrant index (1–4) within that field. ## Provenance The underlying smartphone-microscopy imagery and parasite-coordinate annotations are from: > Morais MCC, Silva D, Milagre MM, de Oliveira MT, Pereira T, Silva JS, et al. > *Automatic detection of the parasite* Trypanosoma cruzi *in blood smears using a machine learning approach applied to mobile phone images.* > PeerJ. 2022;10:e13470. doi:10.7717/peerj.13470 Raw fields and per-field parasite annotations are available through the Morais *et al.* PeerJ Supplemental Information and are **not redeposited here**. This dataset is a derivative work that applies the following transforms: 1. Each 2,448 × 3,264 (or 3,456 × 4,608) field is quartered into four spatial quadrants. 2. Quadrants are standardized to 1,224 × 1,632 pixel lossless PNG. 3. White balance correction (OpenCV `xphoto` with white point estimated from the central 50 % of the image). 4. Each quadrant is labeled positive if at least one Morais (x, y) parasite centroid falls within it, negative otherwise. 5. Slide-level partition at field 0600 (Morais's own train/test boundary): fields 0001–0519 → train, 0520–0599 → validation, 0600–0704 → test. 6. Train split only: flip-balancing of the minority (negative) class using vertical/horizontal/combined flips, followed by a Gaussian-noise-augmented copy (σ = 50 on an 8-bit scale, clipped to [0, 255]) of every balanced-and-flipped training image. The full generation pipeline is open-source at https://github.com/chrislee339/tcruzi-fpn. The script that produced this exact split is `fieldsplit/make_fieldsplit.py`.

提供机构:
Zenodo
创建时间:
2026-04-23
二维码
社区交流群
二维码
科研交流群
商业服务