HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation
Trong-Thuan Nguyen, Pha Nguyen, Jackson Cothren, Alper Yılmaz, Khoa Luu
University of Arkansas at Fayetteville The Ohio State University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Multimodal LLMs have advanced vision-language tasks but still struggle with understanding video scenes. To bridge this gap, Video Scene Graph Generation (VidSGG) has emerged to capture multi-object relationships across video frames. However, prior methods rely on pairwise connections, limiting their ability to handle complex multi-object interactions and reasoning. To this end, we propose the Multimodal Large Language Models (LLMs) on a Scene HyperGraph (HyperGLM), promoting reasoning about multi-way interactions and higher-order relationships. Our approach uniquely integrates entity scene graphs, which capture spatial relationships between objects, with a procedural graph that models their causal transitions, forming a unified HyperGraph. Significantly, HyperGLM enables reasoning by injecting this unified HyperGraph into LLMs. Additionally, we introduce a new Video Scene Graph Reasoning (VSGR) dataset featuring 1.9M frames from third-person, egocentric, and drone views and support five tasks. Empirically, HyperGLM consistently outperforms state-of-the-art methods, effectively modeling and reasoning complex relationships in diverse scenes.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIMultimodal Machine Learning Applications
Video Analysis and Summarization · Human Pose and Action Recognition
参考文献 42
此处列出前 3 条
引用本文 8
按被引量排序,此处列出前 3 条