Table 2.

Statistics of the speech corpora were computed using BERT’s tokenizer to determine the number of tokens in each ground truth sentence

CorpusSplits#Utterances#Hours#Tokens/sentence
Min.Avg.Max.
AISHELL-1Train120,098150.0114.444
Dev14,32618.0314.335
Test7,17610.0314.637
TEDLIUM-2Train92,973212.0126.096
Dev5071.6138.5143
Test1,1552.6126.3132
LibriSpeechTrain281,231960.9136.697
Dev-clean2,7035.4122.1113
Test-clean2,6205.4122.0109
Dev-other2,8645.3119.692
Test-other2,9395.1119.7123

or Create an Account

Close subscription notice
Close access options