Table 1.

Comparison of recent pipeline-based ASR post-processing frameworks in terms of processing stages and information utilization

MethodPipelinesUses N-best listUses extra LMUses speech info
Cross-modal ASR post-processing (Du et al., 2022)Stage 1: Confidence estimation
Stage 2: Error correction via cross-modal fusion (acoustic + text)
Multi-stage LLM correction (Pu et al., 2023)Stage 1: Uncertainty estimation on N-best list
Stage 2: Rule-guided LLM correction
ClozeGER (Hu et al., 2024)Stage 1: N-best transformation into cloze format
Stage 2: Multimodal LLM generative error correction
MMGER (Mu et al., 2024)Stage 1: Multimodal speech-text encoding
Stage 2: Multi-granularity correction (frame + utterance)
SpeechLLM pseudo-labeling (Prakash et al., 2025)Stage 1: Multi-ASR hypothesis fusion
Stage 2: LLM/SpeechLLM correction and pseudo-label generation
LLM error correction w/ Task-Activating prompting (Yang et al., 2023)Stage 1: Hypothesis rescoring via LLM
Stage 2: Error correction using task-activated prompting
Tag and correct (Ziętkiewicz, 2022)Stage 1: Neural tagger detects erroneous spans
Stage 2: Corrector applies context-aware edits
Pinyin regularization for LLM correction (Tang et al., 2024)Stage 1: Pinyin-based text regularization
Stage 2:LLM refinement for Chinese ASR correction

or Create an Account

Close subscription notice
Close access options