Light Sheet Microscopy Foundation Model Pretraining Dataset
收藏资源简介:
This dataset consists of 3D image patches extracted from light sheet microscopy (LSM) volumes with paired text descriptions. Images are drawn from five sources: an internal dataset from the Wu Lab at Weill Cornell Medicine (New York, NY, USA), and four public datasets: SELMA3D MICCAI challenge, Allen human brain, Allen developing mouse brain, and Allen viral labeling of projection neurons. From each dataset, up to 10 non-overlapping 96×96×96 voxel patches were randomly selected with a minimum foreground threshold of 5%, yielding 1,023 patches in total. The patches capture diverse neuroanatomical structures and imaging conditions across species and brain regions. Text descriptions were authored by a domain expert and paraphrased using an LLM to increase linguistic diversity while preserving semantic fidelity. This dataset was used for pretraining a multimodal LSM foundation model.



