Codec-SUPERB@SLT24 special session将于12月3日下午15:00-18:30在澳门举办,届时SemantiCodec、BNN、Uniaudio、VoiceCraft、Moshi等作者们将给主题演讲。
神经音频编解码器将连续音频转换为离散编码,可以帮助开发大型语音模型,近些年来得到了广泛的关注。本次研讨会旨在促进codec领域学者们之间的沟通与交流,促进codec以及speech LM的发展。

我们报告的部分录影和报告幻灯片会更新在以下两个网站


主题演讲
时间:2024/12/03 15:00 - 18:30

Host

Hungyi Lee(台湾大学):https://speech.ee.ntu.edu.tw/~hylee/index.php

Haibin Wu(微软):https://hbwu-ntu.github.io/


演讲人介绍

1、Academia


Prof. Wenwu Wang

(https://www.surrey.ac.uk/people/wenwu-wang/)

演讲主题:Neural Audio Codecs: Recent Progress and a Case Study with SemantiCodec

摘要:The neural audio codec has attracted increasing interest as a highly effective method for audio compression and representation. By transforming continuous audio into discrete tokens, it facilitates the use of large language modelling (LLM) techniques in audio processing. In this talk, we will report recent progress in neural audio codecs, with a particular focus on SemantiCodec, a new neural audio codec for ultra-low bit rate audio compression and tokenization. SemantiCodec features a dual-encoder architecture: a semantic encoder using a self-supervised AudioMAE, discretized using k-means clustering on extensive audio data, and an acoustic encoder to capture the remaining details. The semantic and acoustic encoder outputs are then used to reconstruct audio via a diffusion-model-based decoder. SemantiCodec presents several advantages over previous codecs, which typically operate at high bitrates, are confined to narrow domains like speech, and lack the semantic information essential for effective language modelling. First, SemantiCodec compresses audio into fewer than 100 tokens per second across various audio types, including speech, general audio, and music, while maintaining high-quality output. Second, it preserves substantially richer semantic information from audio compared to all evaluated codecs. We will illustrate these benefits through benchmarking and conclude by discussing potential directions for future research in this field.

讲者介绍:Wenwu Wang is a Professor in Signal Processing and Machine Learning, University of Surrey, UK. He is also an AI Fellow at the Surrey Institute for People Centred Artificial Intelligence. His current research interests include signal processing, machine learning and perception, artificial intelligence, machine audition (listening), and statistical anomaly detection. He has (co)-authored over 300 papers in these areas. He has been recognized as a (co-)author or (co)-recipient of more than 15 accolades, including the 2022 IEEE Signal Processing Society Young Author Best Paper Award, ICAUS 2021 Best Paper Award, DCASE 2020 and 2023 Judge’s Award, DCASE 2019 and 2020 Reproducible System Award, and LVA/ICA 2018 Best Student Paper Award. He is an Associate Editor (2020-2025) for IEEE/ACM Transactions on Audio Speech and Language Processing, and an Associate Editor (2024-2026) for IEEE Transactions on Multimedia. He was a Senior Area Editor (2019-2023) and Associate Editor (2014-2018) for IEEE Transactions on Signal Processing. He is the elected Chair (2023-2024) of IEEE Signal Processing Society (SPS) Machine Learning for Signal Processing Technical Committee, a Board Member (2023-2024) of IEEE SPS Technical Directions Board, the elected Chair (2025-2027) and Vice Chair (2022-2024) of the EURASIP Technical Area Committee on Acoustic Speech and Music Signal Processing, an elected Member (2021-2026) of the IEEE SPS Signal Processing Theory and Methods Technical Committee. He has been on the organising committee of INTERSPEECH 2022, IEEE ICASSP 2019 & 2024, IEEE MLSP 2013 & 2024, and SSP 2009. He is Technical Program Co-Chair of IEEE MLSP 2025. He has been an invited Keynote or Plenary Speaker on more than 20 international conferences and workshops.


Prof. Minje Kim

(https://minjekim.com/)

演讲主题:Future Directions in Neural Speech Communication Codecs

摘要:Neural speech codecs promise high-quality speech at low bitrates but face challenges like increased model complexity and suboptimal quality. This talk presents two approaches to address these issues: generative de-quantization and personalization. With LaDiffCodec, we propose separating representation learning from information reconstruction. It combines a typical end-to-end codec that learns low-dimensional discrete tokens for compact representation and a latent diffusion model that de-quantizes these tokens into high-dimensional continuous space in a generative fashion. To prevent over-smooth speech, we employ “midway-infilling” during diffusion. Subjective tests show our model improves popular neural speech codecs' performance. Next, we introduce personalized neural speech codecs, where we explore personalizing codecs to specific user groups to reduce complexity and enhance perceptual quality. By learning speaker embeddings with a Siamese network from the LibriSpeech dataset, we cluster speakers based on perceptual similarity. Subjective tests reveal this strategy enables model compression without sacrificing—and even improving—speech quality. By decoupling key tasks and introducing personalization, these approaches address current limitations and pave the way for superior speech quality and efficiency at low bitrates in neural speech codecs.

讲者介绍:

Minje Kim is an associate professor in the Dept. of Computer Science at the University of Illinois at Urbana-Champaign. He is also an Amazon Visiting Academic, working at Amazon Lab126. Before then, he was an associate professor at Indiana University (2016-2023). He earned his Ph.D. in Computer Science at UIUC (2016). He worked as a researcher at ETRI, a national lab in Korea, from 2006 to 2011. He received his Master’s and Bachelor’s degrees in the Dept. of Computer Science and Engineering at POSTECH (Summa Cum Laude) and in the Division of Information and Computer Engineering at Ajou University (with honors) in 2006 and 2004, respectively. During his career as a researcher, he has focused on developing machine learning models for audio signal processing applications. He has been on more than 60 patents as an inventor.


Dongchao Yang

(https://dongchaoyang.top/)

演讲主题:Challenges in Developing Universal Audio Foundation Model

摘要:Building a universal audio foundation model for different audio generation tasks, such as text-to-speech, text-to-audio, singing voice synthesis, voice conversion, and speech dialogue, has attract great interest in the audio community. Audio codec and audio modeling strategies are two key points. In this talk, we review the development of audio codec and different audio modeling strategies in the literature, and show our recent works in audio codec and audio modeling methods. Lastly, we summarize the challenges and potential directions in development universal audio foundation model.

讲者介绍:Dongchao Yang is a second-year Ph.D. student at The Chinese University of Hong Kong, supervised by Prof. Helen Meng. Before that, he received his master's degree from Peking University. His research interests encompass the extensive domain of speech and language intelligence, which includes audio foundation models, multi-modal large language models (MLLMs), text-to-speech synthesis (TTS), cross-modal representation learning, among other related areas. Currently, his work focuses on audio-text foundation models and audio codec models.


2、Industrial


Dr. Shang-Wen Li

(https://swdanielli.github.io/)
标题:VoiceCraft: Zero-Shot Speech Editing and TTS in the Wild
摘要:VoiceCraft leverages Encodec to tokenize speech signals; it employs a Transformer decoder architecture and introduces a token rearrangement procedure that combines causal masking and delayed stacking to enable token generation within an existing sequence. On speech editing tasks, VoiceCraft produces edited speech that is nearly indistinguishable from unedited recordings in terms of naturalness, as evaluated by humans; for zero-shot TTS, our model outperforms prior SOTA models including VALLE and XTTS-v2. The models are evaluated on challenging and realistic datasets, that consist of diverse accents, speaking styles, recording conditions, and background noise and music, and our model performs consistently well compared to other models and real recordings. In particular, for speech editing evaluation, we introduce a high quality, challenging, and realistic dataset named RealEdit. We encourage audience to listen to the demos at this website:https://jasonppy.github.io/VoiceCraft_web/
讲者介绍:Shang-Wen Li is a Research Lead and Manager at Meta’s Fundamental AI Research (FAIR) team, and he worked at Apple Siri, Amazon Alexa and AWS before joining FAIR. He completed his PhD in 2016 at MIT in the Spoken Language Systems group of Computer Science and Artificial Intelligence Laboratory (CSAIL). His recent research is focused on multimodal large language models, multimodal representation learning, and spoken language understanding.



Dr. Neil Zeghidour

(https://lienz.github.io/)
标题:Audio Language Models

摘要:Audio analysis and audio synthesis require modeling long-term, complex phenomena and have historically been tackled in an asymmetric fashion, with specific analysis models that differ from their synthesis counterpart. In this presentation, we will introduce the concept of audiolanguage models, a recent innovation aimed at overcoming these limitations. By discretizing audiosignals using a neural audio codec, we can frame both audio generation and audio comprehension as similar autoregressive sequence-to-sequence tasks, capitalizing on the well-established Transformer architecture commonly used in language modeling. This approach unlocks novel capabilities in areas such as textless speech modeling, zero-shot voice conversion, text-to-music generation and even real-time spoken dialogue. Furthermore, we will illustrate how the integration of analysis and synthesis within a single model enables the creation of versatile audio models capable of handling a wide range of tasks involving audio as inputs or outputs. We will conclude by highlighting the promising prospects offered by these models and discussing the key challenges that lie ahead in their development.

讲者介绍:Founded Kyutai in Paris. Previously Staff Research Scientist at Google DeepMind. Teaching "Algorithms for Speech and Natural Language Processing" at master MVA (École Normale Supérieure). Interested in deep learning for audio understanding, audio synthesis, and signal processing.


接受论文
[1] Jiang, Xiao-Hang, et al. "MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios."arXiv preprint arXiv:2411.00464 (2024).
[2] Guo, Haohan, et al. "Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder."arXiv preprint arXiv:2406.02940 (2024).
[3] Li, Jiaqi, et al. "Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation."arXiv preprint arXiv:2409.04016 (2024).
[4] Shi, Jiatong, et al. "ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech."arXiv preprint arXiv:2409.15897 (2024).
[5] Wu, Haibin, et al. "Codec-SUPERB@ SLT 2024: A lightweight benchmark for neural audio codec models."arXiv preprint arXiv:2409.14085 (2024).