Abstract
We present deep regression, a new approach to solving classical spectrum mapping problems, leveraging upon machine learning and big data paradigms. Based on Kolmogorov’s Representation Theorem (1957), a multivariate scalar function can be expressed exactly as a superposition of a finite number of outer functions with another linear combination of inner functions embedded within. Cybenko (1989) developed a Universal Approximation Theorem showing such a scalar function can be approximated by a superposition of sigmoid functions, inspiring a new wave of neural network algorithms. To make the mapping learnable for practical applications, we cast the classical function approximation problems into a nonlinear regression setting using deep neural networks (DNNs) as mapping functions, such that the DNN parameters can be estimated with deep learning and big data configurations. In this talk, we first develop four new theorems to extend the universal approximation theorems from sigmoid to DNNs, and from vector-to-scalar to vector-to-vector regression. We also show that the generalization loss of regression error in machine learning can be decomposed into three terms, approximation, estimation and optimization errors, such that each of them can be tightly bounded, separately.
Many classical speech processing problems, such as enhancement, source separation and dereverberation, can be formulated as finding mapping functions to transform input to output spectra. Our developed theorems also provide some guidelines for parameter and architecture selections in DNN designs. In a series of experiments for high-dimensional nonlinear regression, we validate our theory in terms of representation and generalization powers in machine learning for speech spectrum mapping. As a result, DNN-transformed speech usually exhibits good quality and clear intelligibility under adverse acoustic conditions. Finally, our proposed deep regression framework was also tested on challenging tasks in CHiME-2, CHiME-4, CHiME-5, CHiME-6, REVERB (CHiME-3) and DIHARD III. Based on the top quality achieved in microphone-array based enhancement, separation and dereverberation, our teams scored the lowest error rates in almost all the above-mentioned open evaluation scenarios.
Biography
Chin-Hui Lee is a professor at School of Electrical and Computer Engineering, Georgia Institute of Technology. Before joining academia in 2001, he had accumulated 20 years of industrial experience ending in Bell Laboratories, Murray Hill, as the Director of the Dialogue Systems Research Department. Dr. Lee is a Fellow of the IEEE and a Fellow of ISCA. He has published over 600 papers and 30 patents, with more than 34,500 citations and an h-index of 90 on Google Scholar. He received numerous awards, including the Bell Labs President’s Gold Award in 1998. He won IEEE Signal Processing Society (SPS) 2006 Technical Achievement Award for “Exceptional Contributions to the Field of Automatic Speech Recognition”. In 2012 he gave an ICASSP plenary talk on the future of automatic speech recognition. In the same year he was awarded the ISCA Medal in Scientific Achievement for “pioneering and seminal contributions to the principles and practice of automatic speech and speaker recognition”. His two pioneering papers on deep regression accumulated over 2890 citations and won a Best Paper Award from IEEE SPS in 2019. Recently in 2025, he was awarded SPS’s Signal Processing Letters Best Paper Award for his contribution to developing theories behind MAE-based deep regression.