FoodDiff: A Collaborative Relationship Perception Framework for Food Image Synthesis Using Diffusion Models
Mengling Xu, Sisi You, Bing-Kun Bao
Nanjing University of Posts and Telecommunications
内容与影响
Food image generation is a typical application of text-to-image (T2I) models. The core difference between food image synthesis and other T2I tasks is that there exist complex collaborative relationships among ingredients, cooking actions, and food images, which determine the appearance of dishes. However, existing food image generation models generally ignore or fail to sufficiently utilize such collaborative relationships, which hinders the model from precisely perceiving the shapes and details of food. Furthermore, the pre-training distribution of T2I models is usually noisy and differs from the user-preferred food feature distributions, resulting in deviations from human aesthetics. To address the above issues, we proposeFoodDiff, a collaborative relationship-aware diffusion model for food image generation, which consists of three key components: (1) To perceive collaborative relationships, we propose a collaborative relation module to extract these relations and inject them into the image generation process. (2) To sufficiently interact with the relationships between recipe semantics and food representations, we propose a recipe fusion fine-tuning module to precisely fuse recipe semantics with visual features and fine-tune the pre-trained model. (3) To make the pre-training feature distribution conform to human preference, we introduce an image reward feedback mechanism to optimize the aesthetics of food images. In addition, we propose a high-quality food dataset named Food-Aesthetic with exquisite plates and elaborate annotations. Extensive experiments and human evaluations show that FoodDiff has superior image aesthetics and semantic consistency.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
计算机 / AIGenerative Adversarial Networks and Image Synthesis
Aesthetic Perception and Analysis · Visual Attention and Saliency Detection
参考文献 28
此处列出前 3 条