CM-AVAE: Cross-Modal Adversarial Variational Autoencoder for Visual-to-Tactile Data Generation
Qiyuan Xi, Fei Wang, Liangze Tao, Hanjing Zhang, Xun Jiang, Juan Wu
Southeast University
内容与影响
Vibration acceleration signals allow humans to perceive the surface characteristics of textures during tool-surface interactions. However, acquiring acceleration signals requires a specialized system, which is relatively expensive. Conversely, visual images are more accessible than acceleration signals, and generative models can convert visual images into vibration acceleration signals. Utilizing generative models to generate vibration acceleration signals from visual data circumvents the need for time-consuming actual measurements. Furthermore, this approach can be applied to robot-related tasks. This paper presents a cross-modal adversarial variational autoencoder (CM-AVAE) for visual-to-tactile data generation. Our model incorporates latent space learning from variational autoencoders (VAEs) into generative adversarial networks (GANs) and maps the generator's decoder feature vectors to the discriminator. In addition, a public dataset is chosen to train the model, and relevant evaluation metrics are established to evaluate the model's generated results. The results generated by the CM-AVAE model show significant improvement in objective experiments compared to the baseline models. Furthermore, subjective experimental outcomes also surpass those of the baseline models. Ablation study shows that CM-AVAE introduces latent space learning and maps the decoder feature vectors in the generator to the discriminator, significantly improving the quality of cross-modal data generation.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
生物医学Cell Image Analysis Techniques
Generative Adversarial Networks and Image Synthesis · Industrial Vision Systems and Defect Detection
参考文献 38
此处列出前 3 条
施引文献 7
按被引量排序,此处列出前 3 条