Figure 5.
A line graph compares W E R across T Q L values for five utterance-level and character-level T Q L settings.The horizontal axis plots T Q L Value from 0.0 to 1.0. The vertical axis plots W E R in per cent. Five series are compared. With U T T T Q L, W E R falls from about 100 per cent at 0.1 to 74 per cent at 0.3, 34 per cent at 0.5, 15 per cent at 0.7, 11 per cent at 0.9, and 11 per cent at 1.0. With C H A R T Q L, W E R is about 100 per cent at 0.1 and 0.3, 99 per cent at 0.5, 71 per cent at 0.7, 11 per cent at 0.9, and 9 per cent at 1.0. With U T T plus C H A R and both varied, W E R is about 102 per cent at 0.1 and 0.3, 101 per cent at 0.5, 74 per cent at 0.7, 12 per cent at 0.9, and 8 per cent at 1.0. With U T T plus C H A R, varying L subscript u t t while L subscript char is fixed, W E R remains near 8 to 9 per cent from 0.1 to 1.0. With U T T plus C H A R, varying L subscript char while L subscript u t t is fixed, W E R is about 104 per cent at 0.1, 100 per cent at 0.3, 99 per cent at 0.5, 73 per cent at 0.7, 12 per cent at 0.9, and 9 per cent at 1.0.

WER as a function of inference time TQL values across five configurations. w/ UTT TQL and w/ CHAR TQL are single-modality models sweeping their respective TQL parameters. For the dual-modality model (w/ UTT+CHAR), three sweep strategies are compared: varying both TQL parameters simultaneously, varying only Lutt with Lchar fixed at 0.99, and varying only Lchar with Lutt fixed at 0.99. The shaded region indicates the high-TQL range where all configurations converge

Source: Authors’ own work

or Create an Account

Close subscription notice
Close access options