Waterbird
收藏资源简介:
Waterbird数据集由达姆施塔特工业大学创建,用于研究机器学习模型在分类任务中的捷径学习问题。该数据集包含各种鸟类图像,旨在区分水鸟和陆鸟。数据集的设计使得模型容易依赖背景特征而非鸟类本身的特征,从而引发捷径学习现象。创建过程中,数据集通过引入背景特征与鸟类标签之间的虚假关联来模拟现实世界中的数据偏差。该数据集主要应用于机器学习模型的鲁棒性和泛化能力研究,旨在解决模型在面对复杂和多变数据环境时的决策偏差问题。
The Waterbird dataset was created by Technische Universität Darmstadt to study the shortcut learning problem of machine learning models in classification tasks. This dataset contains various bird images, with the goal of distinguishing between waterbirds and landbirds. The design of the dataset causes models to easily rely on background features rather than the intrinsic features of the birds themselves, thus inducing the shortcut learning phenomenon. During its creation, the dataset simulates real-world data biases by introducing spurious correlations between background features and bird labels. This dataset is primarily applied to research on the robustness and generalization ability of machine learning models, aiming to address the decision-making bias issues of models when facing complex and dynamic data environments.




