GZ2 tag collection well-sampled, balanced and random shuffled version.
收藏资源简介:
tags_well-sampled.zip There are .txt file collection of well-sampled GZ2 galaxy tags, whose file names are the asset ids of gz2 galaxy images. And the determination of well samples follows the procedure describled in Willett et al. 2013. gz2_sdss_id_tags_well-sampled.csv The .csv catalog of well-sampled galaxies with column descriptions: asset_id: a unique identifier assigned to each galaxy in the dataset, directly related to [asset_id].txt (tag file) and [asset_id].jpg (image file) in the gz2 file system. tags: a set of descriptive morphological tags associated with the galaxy, following the GZ2 question tree to tag these sources, aligned with the txt file. tags_well-sampled_balanced_shuffled.zip Since this sample is highly imbalanced, while some morphologies appear much more frequently than others. To mitigate this imbalance and avoid the model biased toward the more frequent labels, we down-sample those overrepresented labels to 2,000 to better align with the occurrence of rarer tags. Such treatment can avoid overfitting and improve generalizability. Moreover, to facilitate the diffusion model’s ability to learn specific visual attributes, we randomly sample and shuffle the tags during training. This approach ensures that the final dataset is balanced, consistent, and diverse, providing an effective foundation for subsequent model training. Finally, we got this collection of tags in well-sample, balanced and random shuffled version. Each txt file is the corresponding morphogical description of galaxy with the filename as asset id. gz2_sdss_id_tags_well-sampled_balanced_shuffled.csv The .csv catalog of well-sampled, balanced and random shuffled galaxies with column descriptions the same as last catalog. The preprocessed data is used as training dataset of our work -- Can AI Dream of Unseen Galaxies? Conditional Diffusion Model for Galaxy Morphology Augmentation. These files were generated by our research group and are not copied from external repositories such as GitHub or HuggingFace. If you use the dataset for reproduction or training model, please cite our work. Feel free to contact us if you meet any related questions.



