Cross-modal learning for SAR target recognition using optical vision foundation models

AI-generated keywords: Synthetic Aperture Radar

AI-generated Key Points

  • Synthetic Aperture Radar (SAR) is known for its versatile, long-range capabilities and ability to operate in almost all weather conditions.
  • Automatic Target Recognition (ATR) in SAR images faces challenges due to limited labeled data, strong speckle interference, and a significant domain gap with optical imagery.
  • Electro-optical (EO) imagery benefits from extensive datasets, clearer visual structure, and robust foundation models.
  • The study explores using vision foundation models trained on optical data to provide class-level supervision for SAR classification.
  • A cross-modal EO to SAR prototype alignment framework is proposed where a frozen EO encoder constructs class-level optical prototypes without strict EO/SAR pairs.
  • The approach enhances SAR classification accuracy compared to other baselines and offers clearer class separation within the trained SAR embedding space.
Also access our AI generated: Comprehensive summary, Lay summary, Blog-like article; or ask questions about this paper to our AI assistant.

Authors: Lucas Hirsch, James R. Hopgood, Javid Khan, Yoann Altmann, Mike E. Davies

Accepted for presentation at SPIE Sensors + Imaging 2026
License: CC BY 4.0

Abstract: Synthetic Aperture Radar (SAR) is an important modality in a wide range of imaging applications due to its versatile, long range and near all weather operating capabilities. However, Automatic Target Recognition (ATR) remains a challenging problem due to limited labelled data, the strong speckle in SAR images and the significant domain gap between SAR and more abundant optical imagery. In contrast, electro-optical (EO) imagery benefits from massive datasets, clearer visual structure and powerful foundation models. In this work, we investigate how vision foundation models trained on optical data can provide class level supervision for SAR classification. We propose a cross-modal EO to SAR prototype alignment framework in which a frozen EO encoder, based on a DINOv3 vision foundation model, is used to construct class level optical prototypes without requiring strict EO/SAR pairs. A SAR model is then trained to classify SAR images while aligning its embeddings to the corresponding EO class prototype. At inference time, the SAR model operates independently, without access to optical imagery. We evaluate our approach on the UNICORNv2 dataset, an EO and SAR dataset of civilian vehicles with heavily speckled images and severe class imbalance. EO prototype alignment improves SAR classification accuracy over frozen DINOv3, SAR only finetuning and unpaired distribution alignment baselines, and t-SNE visualizations provide qualitative evidence of clearer separation among classes in the trained SAR embedding space. These results suggest that optical vision foundation models, despite being trained on visible spectrum imagery, provide transferable information for SAR image classification, offering a practical method for using large scale pretrained vision foundation models across challenging sensing modalities.

Submitted to arXiv on 07 Sep. 2026

Ask questions about this paper to our AI assistant

You can also chat with multiple papers at once here.

AI assistant instructions?

Results of the summarizing process for the arXiv paper: 2609.07753v1

, , , , Synthetic Aperture Radar (SAR) is a crucial imaging modality known for its versatile, long-range capabilities and ability to operate in almost all weather conditions. However, Automatic Target Recognition (ATR) in SAR images poses challenges due to limited labeled data, strong speckle interference, and a significant domain gap between SAR and more abundant optical imagery. In contrast, electro-optical (EO) imagery benefits from extensive datasets, clearer visual structure, and robust foundation models. This study explores the potential of using vision foundation models trained on optical data to provide class-level supervision for SAR classification. The researchers propose a cross-modal EO to SAR prototype alignment framework where a frozen EO encoder, based on a DINOv3 vision foundation model, constructs class-level optical prototypes without strict EO/SAR pairs. Subsequently, a SAR model is trained to classify SAR images while aligning its embeddings with the corresponding EO class prototype. Importantly, at inference time, the SAR model operates independently without access to optical imagery. The effectiveness of this approach is evaluated on the UNICORNv2 dataset—a collection of civilian vehicle EO and SAR images characterized by heavy speckling and severe class imbalance. The results demonstrate that EO prototype alignment enhances SAR classification accuracy compared to frozen DINOv3, SAR-only fine-tuning, and unpaired distribution alignment baselines. Additionally, qualitative evidence from t-SNE visualizations indicates clearer class separation within the trained SAR embedding space. Overall,<kgd> this study suggests that optical vision foundation models—despite being trained on visible spectrum imagery—offer transferable information for improving SAR image classification.</kgd> This novel approach presents a practical method for leveraging large-scale pretrained vision foundation models across challenging sensing modalities like SAR imaging.
Created on 10 Sep. 2026

Assess the quality of the AI-generated content by voting

Score: 0

Why do we need votes?

Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.

Similar papers summarized with our AI tools

Navigate through even more similar papers through a

tree representation

Look for similar papers (in beta version)

By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.

Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.