Speech has become the preferred method of interaction for human-machine interfaces and internet of things (IoT) control; however, traditional air-conduction microphones are susceptible to environmental noise interference and lack sufficient security. To address these challenges, this paper proposes a wearable speech recognition system based on a piezo-micromechanical ultrasonic transducer (PMUT) array and a deep residual network (ResNet-50).
At the hardware level, a PMUT array with a cross-equidistant arrangement based on scandium-doped aluminum nitride was fabricated using standard MEMS processes. The compact layout achieved a 45.45% increase in the fill factor, significantly enhancing reception sensitivity. At the algorithmic level, the speech recognition system directly captures faint laryngeal vibration signals; it uses the short-time Fourier transform to map one-dimensional time-domain signals into high-resolution two-dimensional time-frequency spectra, and uses the residual shortcut mechanism of ResNet-50 to eliminate the vanishing gradient problem in deep networks.
Multitask experimental results demonstrate that the system achieved a high classification accuracy of 93.14% ± 0.76% across seven core phoneme classification tasks covering human vocal fold vibrations.
This work provides a reliable technological foundation for future highly robust human–computer interaction and IoT biometric security locks.
