Chin-Hui Lee
Academia Distinguished Lecture

Prof. Chin-Hui Lee

School of Electrical and Computer Engineering, Georgia Institute of Technology
Title A Journey from Classical Universal Approximation to Deep Regression

Abstract

We present deep regression, a new approach to solving classical spectrum mapping problems, leveraging upon machine learning and big data paradigms. Based on Kolmogorov’s Representation Theorem (1957), a multivariate scalar function can be expressed exactly as a superposition of a finite number of outer functions with another linear combination of inner functions embedded within. Cybenko (1989) developed a Universal Approximation Theorem showing such a scalar function can be approximated by a superposition of sigmoid functions, inspiring a new wave of neural network algorithms. To make the mapping learnable for practical applications, we cast the classical function approximation problems into a nonlinear regression setting using deep neural networks (DNNs) as mapping functions, such that the DNN parameters can be estimated with deep learning and big data configurations. In this talk, we first develop four new theorems to extend the universal approximation theorems from sigmoid to DNNs, and from vector-to-scalar to vector-to-vector regression. We also show that the generalization loss of regression error in machine learning can be decomposed into three terms, approximation, estimation and optimization errors, such that each of them can be tightly bounded, separately.

Many classical speech processing problems, such as enhancement, source separation and dereverberation, can be formulated as finding mapping functions to transform input to output spectra. Our developed theorems also provide some guidelines for parameter and architecture selections in DNN designs. In a series of experiments for high-dimensional nonlinear regression, we validate our theory in terms of representation and generalization powers in machine learning for speech spectrum mapping. As a result, DNN-transformed speech usually exhibits good quality and clear intelligibility under adverse acoustic conditions. Finally, our proposed deep regression framework was also tested on challenging tasks in CHiME-2, CHiME-4, CHiME-5, CHiME-6, REVERB (CHiME-3) and DIHARD III. Based on the top quality achieved in microphone-array based enhancement, separation and dereverberation, our teams scored the lowest error rates in almost all the above-mentioned open evaluation scenarios.

Biography

Chin-Hui Lee is a professor at School of Electrical and Computer Engineering, Georgia Institute of Technology. Before joining academia in 2001, he had accumulated 20 years of industrial experience ending in Bell Laboratories, Murray Hill, as the Director of the Dialogue Systems Research Department. Dr. Lee is a Fellow of the IEEE and a Fellow of ISCA. He has published over 600 papers and 30 patents, with more than 34,500 citations and an h-index of 90 on Google Scholar. He received numerous awards, including the Bell Labs President’s Gold Award in 1998. He won IEEE Signal Processing Society (SPS) 2006 Technical Achievement Award for “Exceptional Contributions to the Field of Automatic Speech Recognition”. In 2012 he gave an ICASSP plenary talk on the future of automatic speech recognition. In the same year he was awarded the ISCA Medal in Scientific Achievement for “pioneering and seminal contributions to the principles and practice of automatic speech and speaker recognition”. His two pioneering papers on deep regression accumulated over 2890 citations and won a Best Paper Award from IEEE SPS in 2019. Recently in 2025, he was awarded SPS’s Signal Processing Letters Best Paper Award for his contribution to developing theories behind MAE-based deep regression.

Khein-Seng Pua
Distinguished Lecture

Mr. Khein-Seng Pua (潘健成)

Co-Founder and CEO, Phison Electronics Corporation
Title From Cloud-AI to Edge AI Deployment

Abstract

Generative AI is rapidly transforming the way businesses operate. However, the real challenge for enterprises is no longer simply how to use AI, but how to balance data security, cost, computing resources, and AI performance when deploying AI at scale. This presentation will share Phison’s practical experience in deploying the Phison AI Data Platform internally, including the key differences between cloud-based and on-premises AI and how enterprises can determine the right approach for different use cases. Through Phison’s own deployment experience, the presentation will explore how companies can turn proprietary enterprise data into valuable AI assets, build enterprise-specific AI applications, and progressively adopt Generative AI and AI Agents. Finally, Phison will share its perspective on the future of enterprise AI and the opportunities for Malaysian businesses to accelerate AI adoption through on-premises AI and a growing AI ecosystem.

Biography

Dato K.S. Pua is the co-Founder and Chief Executive Officer of Phison Electronics, a global leader in NAND flash controllers and storage solutions. Born in Sekinchan, Selangor, Malaysia, Mr. Pua moved to Taiwan at the age of 19 to pursue his education at National Chiao Tung University (now National Yang Ming Chiao Tung University).

In 2000, he co-founded Phison with four university classmates and pioneered the development of the world’s first single-chip USB flash drive controller, helping establish a new era of portable storage.

Under his leadership, Phison has grown from a controller IC design company into a global technology platform integrating NAND controllers, storage solutions, enterprise SSDs, and AI computing technologies. The company is increasingly focused on enterprise SSDs (eSSD), customized storage solutions, and AI infrastructure, including its proprietary aiDAPTIV platform for on-premises and edge AI applications.

With more than 25 years of experience in the semiconductor and storage industries, Mr. Pua continues to lead Phison’s global expansion and innovation while strengthening its technology and talent ecosystem in Malaysia.

Sakriani Sakti
Keynote Speech

Prof. Sakriani Sakti

Nara Institute of Science and Technology (NAIST), Japan
Title Learning Speech Like Infants: Rethinking Data-Hungry AI for Inclusive Speech Technology

Abstract

Modern speech AI has achieved remarkable progress, but its success depends heavily on massive datasets, extensive annotation, and substantial hidden human labor. As a result, the benefits of today’s speech technologies remain concentrated in a small number of well-documented languages, leaving many of the world’s languages—especially those spoken by indigenous and under-resourced communities—without effective technological support. This raises an important question: can we rethink how machines learn speech so that language technologies become more inclusive and accessible to the linguistic diversity of humanity?

In contrast, human infants acquire language from limited, noisy, and unstructured input through continuous listening, speaking, and interaction with their environment. Inspired by these learning processes, this talk explores alternative paradigms for speech technology that move beyond data-intensive approaches, including machine speech chain frameworks, zero-resource speech processing, and textless speech-to-speech translation. I will also discuss broader efforts within the research community to support under-resourced languages and the importance of building collaborations with the communities whose languages we aim to serve. By combining insights from human language acquisition with more inclusive research practices, we may move toward speech technologies that better reflect and support the diversity of human language.

Biography

Sakriani Sakti is Head of the Human-AI Interaction (HAI) Research Laboratory, Head of the International Collaboration Division of Information Science, and Full Professor at the Nara Institute of Science and Technology (NAIST), Japan. She is also a Visiting Research Scientist at the RIKEN Center for Advanced Intelligence Project (RIKEN AIP), an Adjunct Professor at the University of Indonesia, and serves as NAIST Assistant President for International Affairs.

Her research spans speech and language technologies and human-AI interaction, with a particular focus on multilingual and under-resourced languages, including speech recognition and synthesis, speech translation, social-affective dialogue systems, and cognitive communication. She has contributed extensively to multilingual speech and language technologies through numerous international collaborative initiatives, including the Asian Pacific Telecommunity Project and speech-to-speech translation projects such as A-STAR and U-STAR.

She serves as a Board Member of the International Speech Communication Association (ISCA) and the ELRA Language Resources Association (ELRA), and as a Member of the Education Board of the IEEE Signal Processing Society. She currently serves as Convener of Oriental-COCOSDA, representing 18 countries and regions across Asia in the area of spoken language resources and technologies. She played a key role in establishing the ELRA–ISCA Special Interest Group on Under-resourced Languages (SIGUL) and has served as its Chair since 2021. In collaboration with UNESCO and ELRA, she served as General Chair of the Language Technologies for All (LT4All) conferences in 2019 and 2025, promoting linguistic diversity, multilingualism, and inclusive language technologies worldwide. She has also contributed to the organization of major international conferences, including serving as Program Chair of LREC-COLING 2024 and AACL 2025, and currently serves as Program Chair for IEEE SLT 2026.

Yu Tsao
Keynote Speech

Prof. Yu Tsao

Research Center for Information Technology Innovation, Academia Sinica
Title Advancing Assistive Oral Communication Technologies through Artificial Intelligence

Abstract

This presentation provides an overview of AI-driven assistive oral communication technologies, encompassing both assistive speaking and assistive hearing domains. The first part focuses on assistive speaking technologies, highlighting intelligent diagnostic and enhancement frameworks for speech disorders. It introduces machine learning approaches for pathological speech classification, severity assessment, and targeted enhancement for conditions such as dysarthria, post-surgical speech impairment, and electrolaryngeal speech. The second part addresses assistive hearing, presenting recent advances in AI-based diagnostic and signal processing techniques for hearing disorders. Representative applications include automated detection of otitis media with effusion, as well as AI-driven speech generation and objective quality assessment methods for hearing aids and cochlear implants. By integrating speech enhancement, assessment, and generation within a unified AI framework, this presentation demonstrates the potential of neural-based technologies to enhance communication effectiveness and accessibility, while underscoring the importance of interdisciplinary research in advancing next-generation, human-centered assistive systems.

Biography

Yu Tsao (Senior Member, IEEE) received the B.S. and M.S. degrees in Electrical Engineering from National Taiwan University, Taipei, Taiwan, in 1999 and 2001, respectively, and the Ph.D. degree in Electrical and Computer Engineering from the Georgia Institute of Technology, Atlanta, GA, USA, in 2008. From 2009 to 2011, he was a Researcher at the National Institute of Information and Communications Technology (NICT), Tokyo, Japan, where he conducted research and product development in multilingual speech-to-speech translation systems, focusing on automatic speech recognition. He is currently a Research Fellow (Professor) and the Deputy Director at the Research Center for Information Technology Innovation, Academia Sinica, Taipei, Taiwan. He also holds a joint appointment as a Professor in the Department of Electrical Engineering at Chung Yuan Christian University, Taoyuan, Taiwan. His research interests include assistive oral communication technologies, audio coding, and bio-signal processing. He serves as an Associate Editor for IEEE Transactions on Consumer Electronics and IEEE Signal Processing Letters. He received the Outstanding Research Award from Taiwan’s National Science and Technology Council (NSTC), the 2025 IEEE Chester W. Sall Memorial Award, and served as the corresponding author of a paper that won the 2021 IEEE Signal Processing Society Young Author Best Paper Award.

Zhizheng Wu
Keynote Speech

Prof. Zhizheng Wu

The Chinese University of Hong Kong, Shenzhen
Title Controllable and Unified Modeling for Speech Generation

Abstract

This talk explores controllable and unified modeling for speech generation, an emerging paradigm for building general-purpose speech generation systems. I will begin by discussing recent advances in speech representation, including neural codecs, discrete speech tokens, and self-supervised features. I will then introduce unified modeling approaches that integrate tasks such as text-to-speech, voice conversion, singing voice synthesis, and speech enhancement within a shared framework. Finally, I will highlight reinforcement learning for human preference alignment, showing how preference modeling can improve the naturalness, expressiveness, and overall quality of generated speech. The talk will conclude with recent progress and future directions toward controllable, expressive, and human-aligned speech generation.

Biography

Zhizheng Wu is an Associate Professor at The Chinese University of Hong Kong, Shenzhen. He received his Ph.D. from Nanyang Technological University and has held positions at Meta (formerly Facebook), Apple, the University of Edinburgh, and Microsoft Research Asia. Professor Wu has initiated several open-source projects, including Merlin, Amphion, and Emilia. He also initiated and organized the first ASVspoof Challenge and the first Voice Conversion Challenge, and served as an organizer of the Blizzard Challenge 2019. He currently serves on the editorial boards of IEEE/ACM Transactions on Audio, Speech, and Language Processing and IEEE Signal Processing Letters. He was the General Chair of the IEEE Spoken Language Technology Workshop 2024 and will serve as the General Chair of the IEEE Conference on Artificial Intelligence 2027.