Exploring Intra- and Inter-Video Relation for Surgical Semantic Scene Segmentation
Yueming Jin, Yang Yu, Cheng Chen, Zixu Zhao, Pheng‐Ann Heng, Danail Stoyanov
Wellcome / EPSRC Centre for Interventional and Surgical Sciences University College London Chinese University of Hong Kong Shenzhen Institutes of Advanced Technology
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Automatic surgical scene segmentation is fundamental for facilitating cognitive intelligence in the modern operating theatre. Previous works rely on conventional aggregation modules (e.g., dilated convolution, convolutional LSTM), which only make use of the local context. In this paper, we propose a novel framework STswinCL that explores the complementary intra- and inter-video relations to boost segmentation performance, by progressively capturing the global context. We firstly develop a hierarchy Transformer to capture intra-video relation that includes richer spatial and temporal cues from neighbor pixels and previous frames. A joint space-time window shift scheme is proposed to efficiently aggregate these two cues into each pixel embedding. Then, we explore inter-video relation via pixel-to-pixel contrastive learning, which well structures the global embedding space. A multi-source contrast training objective is developed to group the pixel embeddings across videos with the ground-truth guidance, which is crucial for learning the global property of the whole data. We extensively validate our approach on two public surgical video benchmarks, including EndoVis18 Challenge and CaDIS dataset. Experimental results demonstrate the promising performance of our method, which consistently exceeds previous state-of-the-art approaches. Code is available at https://github.com/YuemingJin/STswinCL.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AIDomain Adaptation and Few-Shot Learning
Medical Image Segmentation Techniques · Surgical Simulation and Training
参考文献 68
此处列出前 3 条
引用本文 44
按被引量排序,此处列出前 3 条