The framework contains a Feature Extraction Module and a Self-support matching strategy. The support input is an aircraft photograph. Its features pass through a residual network, EfficientNet version 2 and a Contrastive Language-Image Pretraining visual encoder. A weighted mask feature is combined with a residual network feature and processed by weighted global average pooling. The resulting support information is merged with features processed through a convolution layer, rectified linear unit and dropout to form P subscript s. The query input is another aircraft photograph. Its features pass through a weight-shared residual network and EfficientNet version 2 to form F subscript q. Cosine similarity between P subscript s and F subscript q produces the initial query masks M tilde subscript q comma b and M tilde subscript q comma f. These masks enter the self-support prototype generation stage through A S B P and S S F P. The generated self-support prototypes are P star subscript q comma b and P star subscript q comma f. They are combined with P subscript s to form P star subscript s. A final cosine similarity operation compares P star subscript s with the query feature pathway and produces the segmented aircraft output.Proposed model architecture. The left section illustrates the feature extraction module, which we developed in this research. The module consists of three key components: the shared multi-branch network, weighted mask features and enhanced feature fusion
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.