Comparison of recent pipeline-based ASR post-processing frameworks in terms of processing stages and information utilization
| Method | Pipelines | Uses N-best list | Uses extra LM | Uses speech info |
|---|---|---|---|---|
| Cross-modal ASR post-processing (Du et al., 2022) | Stage 1: Confidence estimation | ✘ | ✓ | ✓ |
| Stage 2: Error correction via cross-modal fusion (acoustic + text) | ||||
| Multi-stage LLM correction (Pu et al., 2023) | Stage 1: Uncertainty estimation on N-best list | ✓ | ✓ | ✘ |
| Stage 2: Rule-guided LLM correction | ||||
| ClozeGER (Hu et al., 2024) | Stage 1: N-best transformation into cloze format | ✓ | ✓ | ✓ |
| Stage 2: Multimodal LLM generative error correction | ||||
| MMGER (Mu et al., 2024) | Stage 1: Multimodal speech-text encoding | ✘ | ✓ | ✓ |
| Stage 2: Multi-granularity correction (frame + utterance) | ||||
| SpeechLLM pseudo-labeling (Prakash et al., 2025) | Stage 1: Multi-ASR hypothesis fusion | ✓ | ✓ | ✓ |
| Stage 2: LLM/SpeechLLM correction and pseudo-label generation | ||||
| LLM error correction w/ Task-Activating prompting (Yang et al., 2023) | Stage 1: Hypothesis rescoring via LLM | ✓ | ✓ | ✘ |
| Stage 2: Error correction using task-activated prompting | ||||
| Tag and correct (Ziętkiewicz, 2022) | Stage 1: Neural tagger detects erroneous spans | ✘ | ✓ | ✘ |
| Stage 2: Corrector applies context-aware edits | ||||
| Pinyin regularization for LLM correction (Tang et al., 2024) | Stage 1: Pinyin-based text regularization | ✘ | ✓ | ✘ |
| Stage 2:LLM refinement for Chinese ASR correction |
| Method | Pipelines | Uses | Uses extra | Uses speech info |
|---|---|---|---|---|
| Cross-modal | ✘ | ✓ | ✓ | |
| Multi-stage | ✓ | ✓ | ✘ | |
| ClozeGER ( | ✓ | ✓ | ✓ | |
| ✘ | ✓ | ✓ | ||
| SpeechLLM pseudo-labeling ( | ✓ | ✓ | ✓ | |
| ✓ | ✓ | ✘ | ||
| Tag and correct ( | ✘ | ✓ | ✘ | |
| Pinyin regularization for | ✘ | ✓ | ✘ | |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.