Figure 3.
A feature fusion framework combines residual network, Efficient Net version 2 and Contrastive Language-Image Pre-training visual encoder features with a weighted mask feature.The input I is an aircraft photograph processed by three branches: a residual network, Efficient Net version 2 and a Contrastive Language-Image Pretraining visual encoder. The residual network produces F 1. The mask M is applied to F 1, and the weighted result passes through weighted global average pooling. The mask area is also calculated. The pooled feature and area information pass through normalisation to produce F 1 star. Efficient Net version 2 produces F 2, while the Contrastive Language-Image Pretraining visual encoder produces F 3. F 2 and F 3 enter a feature fusion block and produce F 2 plus 3 star. In the enhanced feature fusion stage, F 1 star and F 2 plus 3 star are combined. A further residual connection from F 1 is added through the indicated fusion path. The final fused support feature is labelled F subscript s.

Proposed weighted mask feature: the weighted mask feature emphasizes the region of interest by enhancing relevant features and suppressing background noise. The enhanced feature fusion module combines features from multiple backbones and uses intra-domain residual connections (shown as red dashed lines) to improve feature alignment

or Create an Account

Close Modal
Close Modal