Six distinct token mixers benchmarked in this paper, each with their unique property for probing the ATEM framework
| Token mixer | Category | Key property |
|---|---|---|
| Pooling-Mixer | Parameter-free baseline | Establishes the lower bound for spatial context without learnable parameters |
| MLP-Mixer | Fixed topology | Global connectivity; lacks 2D geometric priors, prone to oversmoothing |
| CNN-Mixer | Fixed topology | Enforces strict spatial locality and translation equivariance |
| Attn-Conv Hybrid | Hybrid mechanism | Applies global attention before local feature extraction (sub-optimal) |
| Conv-Attn Hybrid | Hybrid mechanism | Applies local feature priming before global attention (high fidelity) |
| Swin-Mixer | Efficient attention | Localized, shifting windowed attention maintaining an complexity |
| Token mixer | Category | Key property |
|---|---|---|
| Pooling-Mixer | Parameter-free baseline | Establishes the lower bound for spatial context without learnable parameters |
| MLP-Mixer | Fixed topology | Global connectivity; lacks 2D geometric priors, prone to oversmoothing |
| CNN-Mixer | Fixed topology | Enforces strict spatial locality and translation equivariance |
| Attn-Conv Hybrid | Hybrid mechanism | Applies global attention before local feature extraction (sub-optimal) |
| Conv-Attn Hybrid | Hybrid mechanism | Applies local feature priming before global attention (high fidelity) |
| Swin-Mixer | Efficient attention | Localized, shifting windowed attention maintaining an |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.