The process begins with an audio waveform. The same audio input is sent to System 1, System 2 and System 3. Each system produces Transcript 1, Transcript 2 and Transcript 3, respectively. The three transcripts enter an alignment module, which aligns their corresponding content. The aligned outputs pass to a voting module. The voting result produces the best-scoring transcript.