Interpretable Alzheimer’s Disease Detection Via Multi-Scale Fusion of Disentangled Speech Features
Yuehua Chen, Xiaokang Liu, Rongfeng Su, Lan Wang, Nan Yan
Southern University of Science and Technology Chinese Academy of Sciences Shenzhen Institutes of Advanced Technology
内容与影响
Spontaneous speech has emerged as a promising biomarker for the non-invasive detection of Alzheimer’s disease (AD). Existing approaches rely on theory-based acoustic features, which are interpretable but may incompletely capture AD-related speech patterns, and on high-dimensional deep learning representations, which are expressive less interpretable. To address these limitations, this study proposes an interpretable AD detection framework based on multi-scale fusion of disentangled speech representations. A neural audio codec is employed to decompose the speech signal into three interpretable attributes: content, prosody, and timbre. These representations are integrated with sentence-level and global linguistic embeddings through a Graph Attention Network (GAT), enabling the modeling of complex interdependencies across multiple scales. Experimental results demonstrate that our proposed method achieves performance comparable to state-of-the-art models on cross-cultural and cross-linguistic datasets, attaining accuracies of 89.6%, 85.9%, and 95.4% on the ADReSS, ADReSSo, and a Chinese dataset, respectively. Feature analysis further indicates that timbre provides a discriminative signal comparable to linguistic features, suggesting its potential as a biomarker. The results highlight that fusing disentangled, multi-scale speech representations can improve both the performance and interpretability of automated AD detection systems.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
计算机 / AISpeech Recognition and Synthesis
Machine Learning in Healthcare · COVID-19 diagnosis using AI
参考文献 18
此处列出前 3 条