Statistics of the speech corpora were computed using BERT’s tokenizer to determine the number of tokens in each ground truth sentence
| Corpus | Splits | #Utterances | #Hours | #Tokens/sentence | ||
|---|---|---|---|---|---|---|
| Min. | Avg. | Max. | ||||
| AISHELL-1 | Train | 120,098 | 150.0 | 1 | 14.4 | 44 |
| Dev | 14,326 | 18.0 | 3 | 14.3 | 35 | |
| Test | 7,176 | 10.0 | 3 | 14.6 | 37 | |
| TEDLIUM-2 | Train | 92,973 | 212.0 | 1 | 26.0 | 96 |
| Dev | 507 | 1.6 | 1 | 38.5 | 143 | |
| Test | 1,155 | 2.6 | 1 | 26.3 | 132 | |
| LibriSpeech | Train | 281,231 | 960.9 | 1 | 36.6 | 97 |
| Dev-clean | 2,703 | 5.4 | 1 | 22.1 | 113 | |
| Test-clean | 2,620 | 5.4 | 1 | 22.0 | 109 | |
| Dev-other | 2,864 | 5.3 | 1 | 19.6 | 92 | |
| Test-other | 2,939 | 5.1 | 1 | 19.7 | 123 | |
| Corpus | Splits | #Utterances | #Hours | #Tokens/sentence | ||
|---|---|---|---|---|---|---|
| Min. | Avg. | Max. | ||||
| AISHELL-1 | Train | 120,098 | 150.0 | 1 | 14.4 | 44 |
| Dev | 14,326 | 18.0 | 3 | 14.3 | 35 | |
| Test | 7,176 | 10.0 | 3 | 14.6 | 37 | |
| TEDLIUM-2 | Train | 92,973 | 212.0 | 1 | 26.0 | 96 |
| Dev | 507 | 1.6 | 1 | 38.5 | 143 | |
| Test | 1,155 | 2.6 | 1 | 26.3 | 132 | |
| LibriSpeech | Train | 281,231 | 960.9 | 1 | 36.6 | 97 |
| Dev-clean | 2,703 | 5.4 | 1 | 22.1 | 113 | |
| Test-clean | 2,620 | 5.4 | 1 | 22.0 | 109 | |
| Dev-other | 2,864 | 5.3 | 1 | 19.6 | 92 | |
| Test-other | 2,939 | 5.1 | 1 | 19.7 | 123 | |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.