Facial Expression Monitoring via Fine-Grained Vision-Language Alignment
Weihong Ren, Yu Gao, Xi’ai Chen, Zhi Jun Han, Zhiyong Wang, Jiaole Wang, Honghai Liu
Harbin Institute of Technology Shenyang Institute of Automation Chinese Academy of Sciences
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
In the fields of health care and clinical monitoring, vision-based Facial Expression Recognition (FER) has achieved significantly progress, but it still faces the challenge of poor generalization ability under unconstrained conditions of occlusions and pose variation. Recently, Vision-Language Model (VLM) has greatly advanced the FER task. However, the existing VLM-based FER methods typically leverage a hard-crafted prompt (e.g., “a photo of [class]”) and only focus on the holistic semantic alignment, which may suffer from modal heterogeneity. In this work, we propose a fine-grained vision-language model via Prompt Masking for FER (PMFER). Specifically, for each expression, we first create fine-grained prompts using facial action units to guide the image encoder to learn discriminative representations. Further, to finely align text prompts and visual action units, we randomly drop a phrase description in the prompts and then predict the dropped phrase by conducting modal cross attention, implicitly promoting fine-grained vision-language alignment. In addition, we also design a modal-adversarial strategy to holistically eliminate the modal difference between visual and textual embeddings in a common latent space. Experimental results demonstrate that our PMFER model outperforms the state-of-the-art methods on several FER benchmarks, especially under the conditions of occlusions and pose variations. Note to Practitioners—Facial expression recognition is very important in health care and clinical monitoring, which provides an useful tool to assess the psychological and physiological conditions of patients. Although FER has made significant progress with the development of deep learning technologies, it still faces problems in the complex environments (e.g., occlusions and pose variations). To address the above issues, we propose a novel FER method in this work based on the recent vision-language model. It takes RGB image and text prompts as input and finally predicts the expression classification. Different from the existing methods, the proposed PMFER can enable fine-grained modal alignment for facial key units. Compared with the state-of-the-art methods on the public datasets, it can achieve better results, especially under the conditions of occlusions and pose variations. Also, we evaluate the proposed method on a real-world pain dataset, and the results demonstrate that PMFER has a good generalization and can be applied to health care.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIFace recognition and analysis
参考文献 48
此处列出前 3 条
引用本文 3
按被引量排序,此处列出前 3 条