Table 4.

Comparison of rate-distortion performance (BD-rate) and computational complexity (kMACs/pixel, parameter size and inference time) for sequential and channel temporal concatenation implemented with ATEM-Hybrid CA and ATEM-Swin. BD-rate values are calculated on the UVG data set relative to the VCT baseline. Encoding and decoding time are per-frame speed, evaluated on 1080p videos with one RTX3090 GPU

ModelBD-rate (⁠↓⁠)kMACs/pix (⁠↓⁠)Parameters(M) (⁠↓⁠)Enc(s) (⁠↓⁠)Dec (s) (⁠↓⁠)
VCT03085.71542.131.08
ATEM-Hybrid CA + SeqCat−16.64%3139.04156.212.141.09
ATEM-Hybrid CA + ChaCat−9.62%2620.29156.211.430.62
ATEM-Swin + SeqCat−3.02%3064.37155.362.091.06
ATEM-Swin + ChaCat−14.86%2542.5155.361.110.46
Source(s): Authors’ own work

or Create an Account

Close subscription notice
Close access options