Self-Supervised Pretraining and Controlled Augmentation Improve Rare Wildlife Recognition in UAV Images
AI-generated Key Points
- Wildlife censuses are important for assessing living conditions and survival risks of wildlife species.
- Traditional manual surveys are being replaced with automatic counts from UAV imagery and deep learning models.
- Supervised models are typically pretrained on large-scale curated datasets before fine-tuning on target imagery.
- The authors propose a self-supervised pretraining methodology using controlled augmentation techniques combined with MoCo and CLD to improve rare wildlife recognition in UAV images.
- Their method outperforms conventional supervised models trained on ImageNet by a significant margin, even when reducing the number of training animals to just 10%.
- This effectively reduces the amount of required annotations while still enabling high-accuracy model training in highly challenging settings.
- The proposed self-supervised pretraining methodology is a promising approach for improving rare wildlife recognition in UAV images and could significantly reduce the amount of required training data while achieving high accuracy models for automated animal censuses using aerial imagery.
Authors: Xiaochen Zheng, Benjamin Kellenberger, Rui Gong, Irena Hajnsek, Devis Tuia
Abstract: Automated animal censuses with aerial imagery are a vital ingredient towards wildlife conservation. Recent models are generally based on deep learning and thus require vast amounts of training data. Due to their scarcity and minuscule size, annotating animals in aerial imagery is a highly tedious process. In this project, we present a methodology to reduce the amount of required training data by resorting to self-supervised pretraining. In detail, we examine a combination of recent contrastive learning methodologies like Momentum Contrast (MoCo) and Cross-Level Instance-Group Discrimination (CLD) to condition our model on the aerial images without the requirement for labels. We show that a combination of MoCo, CLD, and geometric augmentations outperforms conventional models pre-trained on ImageNet by a large margin. Crucially, our method still yields favorable results even if we reduce the number of training animals to just 10%, at which point our best model scores double the recall of the baseline at similar precision. This effectively allows reducing the number of required annotations to a fraction while still being able to train high-accuracy models in such highly challenging settings.
Ask questions about this paper to our AI assistant
You can also chat with multiple papers at once here.
Assess the quality of the AI-generated content by voting
Why do we need votes?
Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.
Similar papers summarized with our AI tools
Navigate through even more similar papers through atree representation
Look for similar papers (in beta version)
By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.
Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.