Performance comparison between baseline models and CREAM variants under the high-error condition. The text correction and text–speech matching modules can be implemented with different settings, while the text rescoring module is implemented using BERT
| Method | High-error scenario | ||||
|---|---|---|---|---|---|
| AISHELL-1 | TEDLIUM-2 | Libri-other | Libri-clean | ||
| Baselines | |||||
| Top-1 | 8.33 | 12.07 | 11.99 | 4.64 | |
| Oracle | 4.74 | 8.55 | 8.14 | 2.43 | |
| BART | 7.43 | 11.77 | 11.26 | 4.50 | |
| FastCorrect | 6.72 | – | – | – | |
| BERT | 6.05 | 10.37 | 9.91 | 3.58 | |
| CREAM | |||||
| Text correction module | Text–speech matching module | ||||
| BART | ASR | 4.91 | 9.48 | 9.12 | 3.28 |
| BART | MATE | 5.05 | 10.41 | 9.43 | 3.93 |
| BART | MIMICA | 5.25 | 11.33 | 10.74 | 4.83 |
| FastCorrect | ASR | 5.01 | – | – | – |
| FastCorrect | MATE | 4.90 | – | – | – |
| FastCorrect | MIMICA | 5.02 | – | – | – |
| BART | ASR+MATE+MIMICA | 4.63 | 9.29 | 8.80 | 3.16 |
| FastCorrect | ASR+MATE+MIMICA | 4.73 | – | – | – |
| Method | High-error scenario | ||||
|---|---|---|---|---|---|
| AISHELL-1 | TEDLIUM-2 | Libri-other | Libri-clean | ||
| Top-1 | 8.33 | 12.07 | 11.99 | 4.64 | |
| Oracle | 4.74 | 8.55 | 8.14 | 2.43 | |
| 7.43 | 11.77 | 11.26 | 4.50 | ||
| FastCorrect | 6.72 | – | – | – | |
| 6.05 | 10.37 | 9.91 | 3.58 | ||
| Text correction module | Text | ||||
| 5.05 | 10.41 | 9.43 | 3.93 | ||
| 5.25 | 11.33 | 10.74 | 4.83 | ||
| FastCorrect | 5.01 | – | – | – | |
| FastCorrect | – | – | – | ||
| FastCorrect | 5.02 | – | – | – | |
| ASR+MATE+MIMICA | 4.63 | 9.29 | 8.80 | 3.16 | |
| FastCorrect | ASR+MATE+MIMICA | 4.73 | – | – | – |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.