Performance comparison between baseline models and CREAM variants under the low-error condition. The text correction and text–speech matching modules can be implemented with different settings, while the text rescoring module is implemented using BERT
| Method | Low-error scenario | ||||
|---|---|---|---|---|---|
| AISHELL-1 | TEDLIUM-2 | Libri-other | Libri-clean | ||
| Baselines | |||||
| Top-1 | 7.22 | 9.72 | 7.21 | 2.58 | |
| Oracle | 4.14 | 6.28 | 4.45 | 1.34 | |
| BART | 7.53 | 10.66 | 7.42 | 3.05 | |
| FastCorrect | 6.81 | – | – | – | |
| BERT | 5.55 | 8.44 | 6.86 | 2.73 | |
| CREAM | |||||
| Text correction modul | Text–speech matching module | ||||
| BART | ASR | 5.30 | 8.43 | 7.21 | 2.89 |
| BART | MATE | 5.70 | 9.63 | 7.39 | 3.48 |
| BART | MIMICA | 5.79 | 9.80 | 7.28 | 2.95 |
| FastCorrect | ASR | 5.09 | – | – | – |
| FastCorrect | MATE | 5.21 | – | – | – |
| FastCorrect | MIMICA | 5.29 | – | – | – |
| BART | ASR+MATE+MIMICA | 4.98 | 8.13 | 6.72 | 2.70 |
| FastCorrect | ASR+MATE+MIMICA | 4.73 | – | – | – |
| Method | Low-error scenario | ||||
|---|---|---|---|---|---|
| AISHELL-1 | TEDLIUM-2 | Libri-other | Libri-clean | ||
| Top-1 | 7.22 | 9.72 | 7.21 | ||
| Oracle | 4.14 | 6.28 | 4.45 | 1.34 | |
| 7.53 | 10.66 | 7.42 | 3.05 | ||
| FastCorrect | 6.81 | – | – | – | |
| 5.55 | 8.44 | 6.86 | 2.73 | ||
| Text correction modul | Text | ||||
| 5.70 | 9.63 | 7.39 | 3.48 | ||
| 5.79 | 9.80 | 7.28 | 2.95 | ||
| FastCorrect | – | – | – | ||
| FastCorrect | 5.21 | – | – | – | |
| FastCorrect | 5.29 | – | – | – | |
| ASR+MATE+MIMICA | 4.98 | 8.13 | 6.72 | 2.70 | |
| FastCorrect | ASR+MATE+MIMICA | 4.73 | – | – | – |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.