Figure 6.
A three-panel architecture compares separate and joint encoders using self-attention, separable C N N blocks, and window-based attention with token concatenation.The Panel A pathway uses two Encoder sub Separate branches. Each branch applies self-attention followed by Separate C N N and repeats this sequence three times. Their outputs undergo sequence-wise concatenation before entering Encoder sub Joint. Encoder sub Joint contains two stacked self-attention blocks and repeats them twice. Panel B again uses two Encoder sub Separate branches. Each branch applies Separate C N N followed by self-attention and repeats this sequence three times. Their outputs undergo sequence-wise concatenation. Encoder sub Joint then applies two self-attention blocks, repeated twice. Panel C uses two Encoder sub Separate branches with window-based attention. The left branch applies W-Self Attention followed by S W-Self Attention and repeats the pair three times. The right branch shows two S W-Self Attention blocks and repeats them three times. Their outputs undergo channel-wise concatenation before entering Encoder sub Joint. The joint encoder applies W-Self Attention followed by S W-Self Attention, with the pair repeated twice.

Schematic diagrams of the three proposed hybrid and efficient token mixer configurations. (a) Attn-Conv Hybrid Mixer: a hybrid design sequencing multi-head self-attention before depthwise convolution. (b) Conv-Attn Hybrid Mixer: the reverse hybrid configuration, prioritizing local convolutional feature extraction before global attention. (c) Swin-Mixer: an efficient attention mechanism pairing windowed attention (W-MSA) with sliding-window attention (SW-MSA). Each conceptual block comprises a pair of these complementary layers; we stack three such blocks to maintain a total depth of six layers, ensuring consistency with the ATEM setup

Source: Authors’ own work

or Create an Account

Close subscription notice
Close access options