All tutorials will be held at Universiti Sains Malaysia (USM) on Tuesday, November 17, 2026.
Tutorial 1
9:00 AM - 10:30 AM
Automated Analysis of Naturalistic Recordings in Early Childhood: Applications, Challenges, and Opportunities
Jialu Li, University of Arizona
Xulin Fan, University of Illinois Urbana-Champaign
Abstract: Naturalistic child-centered audio recordings provide a unique window into early language development, caregiver–child interaction, and developmental or clinical risk. However, their long duration, sparse speech, background noise, speaker overlap, privacy sensitivity, and the high acoustic variability of infant and child vocalizations create challenges that differ substantially from conventional adult speech processing. This tutorial introduces the emerging field of automated analysis of naturalistic long-form recordings centered on children under three years of age. We will review common recording and annotation practices, accessible datasets, and evaluation considerations, then survey core tasks including speaker-type classification, child vocalization classification, child speech and phoneme recognition, infant-cry verification, and language diarization in multilingual settings. We will examine recent advances in self-supervised speech learning for analyzing these recordings at scale, emphasizing their promise and limitations for child-centered audio. The tutorial concludes with open problems in robustness, data scarcity, benchmark design, privacy, and clinically meaningful evaluation across research, healthcare, education, and clinical practice.
Tutorial 2
11:00 AM - 12:30 PM
Speech Enhancement Using Distributed Microphone Arrays
Jie Zhang, USTC
Abstract: Over past a few decades, microphone array technology has matured and been widely applied in human-computer interaction systems such as video conferencing, smart TVs, mobile communications, and hearing aids. However, in real-world noisy environments or distant speech interaction scenarios, the audio capture quality of conventional microphone arrays with fixed configurations is often difficult to guarantee. With the widespread use of wireless smart devices, distributed microphone arrays (or sometimes called wireless acoustic sensor networks) offer new possibilities for improving sound capture quality in complex, open-domain voice interaction systems, providing advantages in array organization, user experience, and wider acoustic field coverage. Recently, distributed microphone arrays have demonstrated strong application potential across many voice interaction tasks in e.g., smart home, classroom, conference, even effectively covering most traditional microphone array applications. This tutorial will focus on summarizing the current sound capture theory and application methods of distributed microphone arrays, including the background, array organization principles, microphone node effectiveness evaluation, and applications to downstream speech tasks. In addition, this tutorial will briefly discuss key challenges and development trends toward practical deployment.
Lunch Break (12:30 PM - 2:00 PM)
Tutorial 3
2:00 PM - 3:30 PM
Securing the Voice: A Tutorial on Audio Deepfake Mitigation
Xuechen Liu, Xi-an Jiaotong-Liverpool University
Abstract: This tutorial traces the technical evolution of voice biometrics security, from basic speaker verification and spoofing detection, to the up-to-date solutions based on Deep Learning and foundational models. The tutorial begins with task definitions, threat models (i.e., synthetic and physical spoofing attacks) and legacy detection frameworks, to a more retrospective overview of the ASVspoof challenge series, spoof-aware speaker verification (SASV), and the APSIPA 2026 RADAR Grand Challenge. Honest critiques will follow up, focusing on generalization in real-world complex scenarios and rapid adaptation. Finally, this tutorial argues that as the boundary between human and machine speech blurs, audio security may move beyond classification toward advanced audio forensics, where the presenter shares some of his personal thoughts. The tutorial targets academic researchers, students, practitioners and newcomers who are interested in speech processing, machine learning, and voice biometrics.
Tutorial 4
4:00 PM - 5:30 PM
Evaluation of Chinese Spoken Language Generation: From TTS to Spoken Dialogue
Yujia Xiao, CUHK
Abstract: This tutorial provides a systematic overview of existing methods and practices for evaluating spoken language generation systems, covering applications such as text-to-speech and spoken dialogue. As speech generation technologies become increasingly natural, expressive, interactive, and context-aware, evaluation must move beyond isolated measures of signal quality or naturalness to consider linguistic correctness, prosodic appropriateness, contextual relevance, interaction quality, robustness, user experience, and trustworthiness. The tutorial will review widely used human and automatic evaluation methods, discuss their strengths and limitations, and organize them according to different evaluation goals and application scenarios. It will also cover practical issues such as evaluation protocol design, diagnostic test set construction, annotation rubrics, subjective and objective metric selection, and failure analysis. While broadly applicable to spoken language generation, the tutorial will include examples and considerations relevant to Chinese spoken language processing, such as language-specific pronunciation, prosody, code-switching, and interaction phenomena.