Figure 1.
A bubble chart compares computation trade-offs among V C T and A T E M variants using k M A C s per pixel and B D-Rate, highlighting differences in computational cost and compression performance.The chart is titled Computation Trade-off. The horizontal axis represents k M A C s per pixel from about 2400 to 4200, and the vertical axis represents B D-Rate in per cent from about minus 20 to 15. V C T is positioned near 3080 k M A C s per pixel with a B D-Rate around 0 per cent. A T E M-Pool appears near 2700 k M A C s per pixel with the highest positive B D-Rate of about 12.5 per cent. A T E M-M L P is near 4010 k M A C s per pixel with a B D-Rate of about 8 per cent. A T E M-Conv is near 3190 k M A C s per pixel with a B D-Rate of about minus 8 per cent. A T E M-Hybrid A C is near 3140 k M A C s per pixel with a B D-Rate around minus 13.5 per cent. A T E M-Hybrid C A is near 3140 k M A C s per pixel with the lowest B D-Rate among the hybrid variants at about minus 16.5 per cent. A T E M-Swin is positioned near 2550 k M A C s per pixel with a B D-Rate around minus 15 per cent, combining relatively low computational cost with a strongly negative B D-Rate. Bubble sizes differ across methods, indicating an additional comparative magnitude, though no separate size scale is shown.

Analysis of the trade-off between rate-distortion performance and computational complexity across varying token mixers. The vertical axis denotes BD-rate (Bjøntegaard, 2001) savings on the UVG data set normalized to the VCT baseline (lower is better), while the horizontal axis represents the computational cost in kMACs per pixel. The bubble size corresponds to the total parameter count of each model. Notably, the ATEM-Swin configuration achieves the most favorable balance between coding efficiency and computational cost among the six tested token mixers

Source: Authors’ own work

or Create an Account

Close subscription notice
Close access options