An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep Learning
Berrak Şişman, Junichi Yamagishi, Simon King, Haizhou Li
Singapore University of Technology and Design University of Edinburgh National University of Singapore
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Speaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic content unchanged. Voice conversion involves multiple speech processing techniques, such as speech analysis, spectral conversion, prosody conversion, speaker characterization, and vocoding. With the recent advances in theory and practice, we are now able to produce human-like voice quality with high speaker similarity. In this article, we provide a comprehensive overview of the state-of-the-art of voice conversion techniques and their performance evaluation methods from the statistical approaches to deep learning, and discuss their promise and limitations. We will also report the recent Voice Conversion Challenges (VCC), the performance of the current state of technology, and provide a summary of the available resources for voice conversion research.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AISpeech Recognition and Synthesis
Speech and Audio Processing · Music and Audio Processing
参考文献 413
此处列出前 3 条
引用本文 360
按被引量排序,此处列出前 3 条