Figure 13.
A quality comparison for H E V C Class B and H E V C Class C shows original crops, five coding methods, and I frame and P frame metrics.The H E V C Class B sequence is shown at 1920 by 1024 pixels. The original scene includes an enlarged crop of a table area. V C T Baseline records I frame and P frame b p p values of 0.499 and 0.169, with P S N R values of 33.15 and 33.47. A T E M-Pool records b p p values of 0.503 and 0.175, with P S N R values of 33.15 and 33.43. A T E M-M L P records b p p values of 0.515 and 0.176, with P S N R values of 33.10 and 33.36. A T E M-Hybrid C A records b p p values of 0.493 and 0.162, with P S N R values of 33.23 and 33.62. A T E M-S w i n records b p p values of 0.481 and 0.155, with P S N R values of 33.19 and 33.58. The H E V C Class C sequence is shown at 832 by 448 pixels. The original scene includes an enlarged crop of product packaging. V C T Baseline records I frame and P frame b p p values of 0.522 and 0.132, with P S N R values of 33.51 and 33.84. A T E M-Pool records b p p values of 0.535 and 0.139, with P S N R values of 33.53 and 33.80. A T E M-M L P records b p p values of 0.536 and 0.136, with P S N R values of 33.52 and 33.79. A T E M-Hybrid C A records b p p values of 0.525 and 0.122, with P S N R values of 33.61 and 34.07. A T E M-S w i n records b p p values of 0.511 and 0.116, with P S N R values of 33.54 and 34.01.

Visual comparison of compression results between the VCT baseline and various token mixer configurations on the HEVC Class B and Class C data sets. All models were evaluated using a fixed Lagrange multiplier of λ=0.01⁠. The reported bits-per-pixel (bpp) and PSNR metrics distinguish between I-frames (Frame 1) and P-frames (Frame 3) within a single Group of Pictures (GoP), highlighting the impact of temporal context on coding efficiency

Source: Authors’ own work

or Create an Account

Close subscription notice
Close access options