Automated location invariant animal detection in camera trap images using publicly available data sources
收藏资源简介:
1. A time-consuming challenge faced by ecologists is the extraction of meaningful data from camera trap images to inform ecological management. Automated object detection solutions are increasingly, however, most are not sufficiently robust to be deployed on a large scale due to lack of location invariance across sites. This prevents optimal use of ecological data and results in significant resource expenditure to annotate and retrain object detectors. 2. In this study, we aimed to (a) assess the value of publicly available image datasets including FlickR and iNaturalist (FiN) when training deep learning models for camera trap object detection (b) develop a for training location invariant object detection models and (c) explore the use of small subsets of camera trap images for optimization training. 3. We collected and annotated 3 datasets of images of striped hyena, rhinoceros and pig, from FiN, and used transfer learning to train 3 object detection models in the task of animal detection. We compared the performance of these models to that of 3 models trained on the Wildlife Conservation Society and Camera CATalogue datasets, when tested on out of sample Snapshot Serengeti datasets. Furthermore, optimized the FiN models via infusion of small subsets of camera trap images to increase robustness for challenging detection cases. 4. In all experiments, the mean Average Precision (mAP) of the FiN models was significantly higher (82.33-88.59%) than that achieved by the models trained only on camera trap datasets (38.5-66.74%). The infusion of camera trap images into FiN training further improved mAP, with increases ranging from 1.78-32.08%. 5. Ecology researchers can use FiN images for training robust, location invariant, out-of-the-box, deep learning object detection solutions for camera trap image processing. This would allow AI technologies to be deployed on a large scale in ecological applications. Datasets and code related to this study are open source and available on Dryad (https://doi.org/10.5061/dryad.1c59zw3tx).
1. 生态学家面临的一项耗时挑战,是从相机陷阱(camera trap)影像中提取有价值的数据以支撑生态管理工作。自动化目标检测方案的应用愈发广泛,但由于缺乏跨站点的位置不变性,多数方案的鲁棒性不足以支撑大规模部署。这不仅阻碍了生态数据的最优利用,还会因标注和重新训练目标检测器而产生大量资源消耗。 2. 本研究旨在达成以下目标:(a) 评估公开影像数据集Flickr与iNaturalist(合称FiN)在训练用于相机陷阱目标检测的深度学习模型时的应用价值;(b) 研发适用于训练位置不变性目标检测模型的方法;(c) 探索使用少量相机陷阱影像子集开展优化训练的可行性。 3. 我们从FiN数据集中收集并标注了斑鬣狗、犀牛与猪三类动物的影像数据集,采用迁移学习训练了3个用于动物检测的目标检测模型。将这些模型的性能与仅在野生动物保护协会(Wildlife Conservation Society)与Camera CATalogue数据集上训练的3个模型进行对比,并在留出的Snapshot Serengeti(快照塞伦盖蒂)数据集上开展测试。此外,我们通过引入少量相机陷阱影像子集对FiN训练得到的模型进行优化,以提升其在复杂检测场景下的鲁棒性。 4. 在所有实验中,FiN训练得到的模型的平均精度均值(mean Average Precision, mAP)显著高于仅在相机陷阱数据集上训练的模型(82.33%~88.59% vs 38.5%~66.74%)。向FiN训练流程中注入相机陷阱影像可进一步提升模型的mAP值,提升幅度介于1.78%~32.08%之间。 5. 生态研究人员可利用FiN数据集的影像,训练用于相机陷阱影像处理的鲁棒性强、具备位置不变性且开箱即用的深度学习目标检测方案。这将推动人工智能技术在生态应用领域实现大规模部署。本研究相关的数据集与代码已开源,可在Dryad平台(https://doi.org/10.5061/dryad.1c59zw3tx)获取。



