Vector Quantized Diffusion Model Based Speech Bandwidth Extension
Yuan Fang, Jing Bai, Jiajie Wang, Xueliang Zhang
Inner Mongolia University Sensetime (China)
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Recent advancements in neural audio codec (NAC) unlock new potential in audio signal processing. Studies have increasingly explored leveraging the latent features of NAC for various speech signal processing tasks. This paper introduces the first approach to speech bandwidth extension (BWE) that utilizes the discrete features obtained from NAC. By restoring high-frequency details within highly compressed discrete tokens, this approach enhances speech intelligibility and naturalness. Based on Vector Quantized Diffusion, the proposed framework combines the strengths of advanced NAC, diffusion models, and Mamba-2 to reconstruct high-frequency speech components. Extensive experiments demonstrate that this method exhibits superior performance across both log-spectral distance and ViSQOL, significantly improving speech quality.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIAdvanced Data Compression Techniques
Speech and Audio Processing
参考文献 22
此处列出前 3 条
引用本文 3
按被引量排序,此处列出前 3 条