标杆管理
增采样
计算机科学
人工智能
特征(语言学)
分割
水准点(测量)
计算机视觉
集合(抽象数据类型)
机器学习
图像分割
特征提取
基础(证据)
模式识别(心理学)
相似性(几何)
编码(集合论)
限制
图像(数学)
人气
特征学习
作者
Volodymyr Havrylov,Haiwen Huang,Dan Zhang,Andreas Geiger
标识
DOI:10.1109/iccvw69036.2025.00028
摘要
Vision Foundation Models (VFMs) are large-scale, pretrained models that serve as general-purpose backbones for various computer vision tasks. As VFMs' popularity grows, there is an increasing interest in understanding their effectiveness for dense prediction tasks. However, VFMs typically produce low-resolution features, limiting their direct applicability in this context. One way to tackle this limitation is by employing a task-agnostic feature upsampling module that refines VFM features resolution. To assess the effectiveness of this approach, we investigate Interactive Segmentation (IS) as a novel benchmark for evaluating feature upsampling methods on VFMs. Due to its inherent multimodal input, consisting of an image and a set of userdefined clicks, as well as its dense mask output, IS creates a challenging environment that demands comprehensive visual scene understanding. Our benchmarking experiments show that selecting appropriate upsampling strategies significantly improves VFM features quality. The code is released at https://github.com/havrylovv/iSegProbe.
科研通智能强力驱动
Strongly Powered by AbleSci AI