Special Session

Beyond Transcript: Affective and Pathological Speech Modeling in the Era of Generative Foundation Models

Traditional speech processing has transitioned from task-specific pipelines to unified LLMs and Foundation Models. However, while these models excel at linguistic content, they often struggle with the "paralinguistic nuance" essential for Affective Computing and Pathological Speech Analysis. This session focuses on the intersection of generative AI and human-centric speech: (1) Modeling Emotion: Moving from discrete labels to nuanced, context-aware affective synthesis and recognition, (2) Clinical Utility: Leveraging LLMs for the early detection of neurodegenerative or mental health disorders (e.g., Alzheimer's, Depression) through vocal biomarkers, (3) Ethics & Alignment: Ensuring that medical/pathological speech applications are safe, private, and bias-free. The motivation is to bridge the gap between "what is said" and "how it is said," fostering a multidisciplinary community at ISCSLP to address these high-stakes applications.

Organizers
  • Mengyue Wu — Shanghai Jiao Tong University
  • Zixing Zhang — Hunan University
  • Kun Qian — Beijing Institute of Technology
  • Zhaojie Luo — Southeast University
  • Ying Shen — Tongji University
  • Wen Wu — Shanghai AI Laboratory

Challenges

CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech

Recent advances in text-to-speech (TTS) have greatly improved speech naturalness, speaker similarity, and controllability. However, most existing controllable TTS systems still rely on explicit user-provided style prompts, making it difficult to automatically determine how a sentence should be spoken in long and complex conversational scenarios. This challenge aims to evaluate whether a system can infer the intended speaking manner from contextual information and generate speech consistent with both the reasoning output and the surrounding scene:

  • Track 1Text-Context-Aware CoT-TTS
  • Track 2Audio-Context-Aware CoT-TTS
Organizers
  • Wei Xue — The Hong Kong University of Science and Technology
  • Junlan Feng — China Mobile
  • Shilei Zhang — Jiutian Artificial Intelligence Technology, China Mobile
  • Yue Wang — China Mobile (Hong Kong) Innovation Research Institute
  • Ruosong Yang — China Mobile (Hong Kong) Innovation Research Institute
  • Bei Liu — The Hong Kong University of Science and Technology
  • Liumeng Xue — Nanjing University
  • Sitong Cheng — The Hong Kong University of Science and Technology
  • Jiahao Pan — The Hong Kong University of Science and Technology
  • Weizhen Bian — The Hong Kong University of Science and Technology
  • Boyi Kang — The Hong Kong University of Science and Technology
  • Bin Long — Hong Kong Generative AI Research & Development Center

Real-World Audio-Visual Speech Enhancement (AVSE) Challenge

Speech enhancement has advanced rapidly as a key front-end technology for intelligent voice interaction, hearing aids, and remote conferencing. Yet purely audio-based methods still face clear performance limits in challenging acoustic conditions, including low SNR, strong reverberation, and overlapping speech from multiple speakers. Visual cues (lip movements, facial information) remain unaffected by acoustic noise, making them a valuable complement to audio. This challenge addresses the gap between academic benchmarks and real-world deployment through two complementary tracks:

  • Track 1Real-World Mixed Scenarios
  • Track 2Visual Degradation
Organizers
  • Kai Li — Tsinghua University
  • Wenze Ren — National Taiwan University
  • Junjie Li — The Hong Kong Polytechnic University
  • Cheng Yu — Ohio State University
  • Peijun Yang — Wuhan University
  • Haibin Wu — Meta
  • Szu-Wei Fu — NVIDIA
  • Wen-Chin Huang — Nagoya University
  • Hsin-Ming Wang — Academia Sinica
  • Xiaolin Hu — Tsinghua University
  • Ming Li — Chinese University of Hong Kong (Shenzhen)
  • Deliang Wang — Chinese University of Hong Kong (Shenzhen)
  • Yu Tsao — Academia Sinica

NVSpeech Challenge: Understanding and Generation of Speech with Non-Verbal Vocalizations

Human speech communication is not limited to linguistic content. In natural conversations, speakers frequently produce non-verbal vocalizations (NVVs), such as laughter, sighs, crying, coughing, breathing, gasps, yawns, and other paralinguistic sounds. These vocalizations convey emotion, attitude, feedback, turn-taking cues, social intent, and sometimes physiological or health-related information. However, most current spoken language processing systems remain largely transcript-centric: ASR systems often ignore or normalize NVVs, while TTS systems usually focus on generating fluent linguistic content without accurately controlling such expressive vocal events. This challenge aims to raise broader attention to NVVs as a crucial yet under-explored component of spoken language communication, with two core tasks:

  • Task 1Understanding: Detect, classify, and temporally localize NVVs in natural speech
  • Task 2Generation: Synthesize speech conditioned on explicit NVV labels or natural-language descriptions
Organizers
  • Xinyuan Qian — University of Science and Technology Beijing
  • Shuai Wang — Nanjing University
  • Jingbin Hu — Northwestern Polytechnical University
  • Yuang Cao — Northwestern Polytechnical University
  • Liumeng Xue — Nanjing University
  • Lei Xie — Northwestern Polytechnical University
  • Wei Xue — The Hong Kong University of Science and Technology
  • Hui Bu — AISHELL
  • Haibin Wu — Meta

Submission Information

Authors are invited to submit original, unpublished papers to the special session and challenges.

  • Regular paper: Up to 4 pages (two-column format), plus up to 2 additional pages for references
  • Long Paper: Up to 8 pages (two-column format), plus up to 2 additional pages for references

Submissions will be reviewed in a double-blind process and must be submitted electronically through the conference submission system. The authors are permitted to post preprints of their work (e.g., on arXiv) at any time.

CMT Submission System: https://cmt3.research.microsoft.com/ISCSLP2026


Submit Your Paper via CMT

Important Dates

  • Paper Submission Deadline: August 3, 2026, at 23:59 (AoE)
  • Paper Acceptance Notification: August 31, 2026 (AoE)
  • Camera-Ready Deadline: September 21, 2026 (AoE)
  • Conference Dates: November 14–17, 2026 (AoE)