Figure 6.
A line graph compares alpha values across training steps for four A P E and P prime T Q L configurations.The horizontal axis shows Training Steps in thousands, from 0 to 90. The vertical axis shows alpha Value from 0.75 to 1.00. All four series begin at about 1.00 at 0 training steps and drop sharply by about 9 thousand steps. With A P E, alpha falls to about 0.78, rises to about 0.85 near 27 thousand steps, then gradually declines to about 0.80 at 90 thousand steps. With P prime subscript U T T T Q L, alpha falls to about 0.80, rises to a peak near 0.88 around 45 thousand steps, then declines gradually to about 0.85 at 90 thousand steps. With P prime subscript C H A R T Q L, alpha falls to about 0.80, rises to about 0.86 around 36 thousand steps, then decreases to about 0.83 at 90 thousand steps. With P prime subscript U T T plus C H A R T Q L, alpha falls to about 0.79, rises to about 0.86 around 36 thousand steps, then declines to about 0.83 at 90 thousand steps.

Evolution of the learned scaling coefficient α during training. TQL configurations converge to higher α (0.8270.855) than APE only (0.804), indicating that explicit quality supervision enables more effective utilization of positional cues. Among TQL configurations, character-level TQL converges to lower α than utterance-level TQL, as per character labels partially fulfill the role of positional encoding

or Create an Account

Close subscription notice
Close access options