Figure 12.
A visual comparison of the original and five coding methods for U V G and M C L-J C V at 1920 by 1080, with enlarged regions and coding metrics.The U V G comparison shows the original portrait and enlarged eye region alongside V C T Baseline, A T E M-Pool, A T E M-M L P, A T E M-Hybrid C A, and A T E M-S w i n results. Metrics are reported for I-frame and P-frame coding. V C T Baseline has b p p values of 0.094 and 0.038, with P S N R values of 34.02 and 34.16. A T E M-Pool has b p p values of 0.093 and 0.039, with P S N R values of 34.00 and 34.12. A T E M-M L P has b p p values of 0.094 and 0.038, with P S N R values of 34.01 and 34.13. A T E M-Hybrid C A has b p p values of 0.090 and 0.034, with P S N R values of 34.02 and 34.17. A T E M-S w i n has b p p values of 0.088 and 0.034, with P S N R values of 34.01 and 34.16. The M C L-J C V comparison shows the original indoor scene and an enlarged region around the woman's face alongside the same five methods. V C T Baseline has b p p values of 0.074 and 0.015, with P S N R values of 38.12 and 38.42. A T E M-Pool has b p p values of 0.075 and 0.016, with P S N R values of 38.09 and 38.37. A T E M-M L P has b p p values of 0.074 and 0.015, with P S N R values of 38.11 and 38.37. A T E M-Hybrid C A has b p p values of 0.071 and 0.012, with P S N R values of 38.62 and 38.41. A T E M-S w i n has b p p values of 0.069 and 0.013, with P S N R values of 38.07 and 38.40.

Visual comparison of compression results between the VCT baseline and various token mixer configurations on the UVG and MCL-JCV data sets. All models were evaluated using a fixed Lagrange multiplier of λ=0.01⁠. The reported bits-per-pixel (bpp) and PSNR metrics distinguish between I-frames (Frame 1) and P-frames (Frame 3) within a single Group of Pictures (GoP), highlighting the impact of temporal context on coding efficiency

Source: Authors’ own work

or Create an Account

Close subscription notice
Close access options