Speech separation with neural audio codecs
I study how the layered representations of neural audio codecs can help separate overlapping voices. My work focuses on preserving the coarse-to-fine structure of residual vector quantization (RVQ) layers.
Speech separation, neural audio codecs, and robust speech recognition.
I study how the layered representations of neural audio codecs can help separate overlapping voices. My work focuses on preserving the coarse-to-fine structure of residual vector quantization (RVQ) layers.
2026
Interspeech 2026 · Accepted
RVQ-Grid explores the structure of residual vector quantization layers for speech separation in neural audio codec representations.
IEICE Technical Report · SP2026-23
Cross-layer recurrent processing preserves the coarse-to-fine structure of RVQ layers for codec-based speech separation.