Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
Ranjan Sapkota, Manoj Karkee
Cornell University
内容与影响
The integration of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional architectures. This in-depth review presents a structured exploration of the state-of-the-art in LVLMs, systematically organized through a three-step research review process. First, we discuss the functioning of vision language models (VLMs) for object detection, describing how these models harness natural language processing (NLP) and computer vision (CV) techniques to revolutionize object detection and localization. We then explain the architectural innovations, training paradigms, and output flexibility of recent LVLMs for object detection, highlighting how they achieve advanced contextual understanding for object detection. The review thoroughly examines the integration of visual and textual information, demonstrating the progress made in object detection using VLMs that facilitate more sophisticated object detection and localization strategies. This review presents visualizations demonstrating LVLMs’ effectiveness in diverse scenarios—including localization and segmentation, compares their real-time performance, adaptability, and complexity to traditional deep learning systems, anticipates LVLMs soon surpassing conventional methods in object detection, highlights the potential of hybrid models combining both approaches’ strengths, and concludes with a discussion on the transformative impact and future roadmap of LVLMs in this field.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
回答优先基于摘要、文献信息与可获取全文;依据不足时会明确说明。
学术脉络
学科主题
计算机 / AIMultimodal Machine Learning Applications
Advanced Image and Video Retrieval Techniques