Cross-Modal Federated Learning among Unimodal Devices
Yongheng Deng, Ningxin He, Xinyi Li, Fan Wu, Yaoxue Zhang, Ju Ren
Tsinghua University Nankai University Central South University
内容与影响
Federated multimodal learning enables decentralized devices with diverse modalities to collaboratively train multimodal models without sharing their raw data. In most existing federated multimodal learning approaches, multimodal data is indispensable. However, in reality, a considerable number of devices can only collect unimodal data, and the data labels are often incomplete. Therefore, this paper proposes FedCMD, a federated multimodal learning approach that enables unimodal devices with missing labels to collaboratively train a multimodal model. To effectively leverage the label-missing samples, FedCMD performs unimodal federated learning first to learn unimodal encoders and make pseudo-labels for the unlabeled samples. Then it calculates and shares the prototypes of various modalities among devices for cross-modal feature alignment. The prototypes are finally served as the complementary of missing modalities to learn a multimodal fusion and classification network. With FedCMD, the learned multimodal network can support both unimodal inputs and multimodal inputs with any missing modalities. Extensive experiments demonstrate the efficacy of FedCMD compared to state-of-the-art baselines.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
计算机 / AISpeech Recognition and Synthesis
User Authentication and Security Systems · Text and Document Classification Technologies
参考文献 43
此处列出前 3 条
施引文献 1
按被引量排序,此处列出前 3 条