On 27 February, the IEEE Signal Processing Society is pleased to welcome Distinguished IEEE Speaker Dr. Tomohiro Nakatani for a guest lecture at Paderborn University. He will present a talk entitled "Enhancing Distant Automatic Speech Recognition via Model-Based Multi-Microphone Front-Ends". The presentation will take place in room L3.204 at 11 am at Paderborn University. Tomohiro Nakatani is a Senior Distinguished Researcher at the Communication Science Laboratories, NTT, Inc, Japan. He was a member of the IEEE Signal Processing Society Audio and Acoustic Signal Processing Technical Committee (2009-2014) and the Speech and Language Processing Technical Committee (2016-2021). In 2021 he was appointed IEEE Fellow.
Abstract:
Distant Automatic Speech Recognition (DASR) refers to the task of recognizing speech captured by farfield microphones. It supports a wide range of applications, including the recognition of natural human conversations in everyday environments. A major challenge in DASR is maintaining high recognition accuracy in the presence of interfering signals such as background noise, reverberation, and overlapping speech.
This talk will provide an overview of modelbased multimicrophone frontend techniques developed to suppress interference in DASR. A key strength of this approach is its ability to decompose signals into individual components using physical and probabilistic signal models, without necessarily requiring prior training. This property enables strong adaptability to unknown and complex environments. Furthermore, when combined with neural network approaches, this framework enables highly accurate frontend processing under adverse conditions.
Through challenging DASR scenarios, the talk will demonstrate how dereverberation, denoising, and source separation frontends can substantially enhance recognition performance.