计算机科学
人工智能
计算机视觉
编码(集合论)
完备性(序理论)
可视化
基础(证据)
建筑
感知
机器学习
图像(数学)
图像处理
下游(制造业)
迭代重建
学习迁移
特征提取
视觉对象识别的认知神经科学
模式识别(心理学)
机器视觉
目标检测
计算模型
源代码
序列(生物学)
作者
Zhaozhi Wang,Yunjie Tian,Zhu Liu,Yaowei Wang,Kun Fu
标识
DOI:10.1109/tip.2026.3652371
摘要
In this study, we introduce EinsPT, an efficient instance-aware pre-training paradigm designed to reduce the transfer gap between vision foundation models and downstream instance-level tasks. Unlike conventional image-level pre-training that relies solely on unlabeled images, EinsPT leverages both image reconstruction and instance annotations to learn representations that are spatially coherent and instance discriminative. To achieve this efficiently, we propose a proxy-foundation architecture that decouples high-resolution and low-resolution learning: the foundation model processes masked low-resolution images for global semantics, while a lightweight proxy model operates on complete high-resolution images to preserve fine-grained details. The two branches are jointly optimized through reconstruction and instance-level prediction losses on fused features. Extensive experiments demonstrate that EinsPT consistently enhances recognition accuracy across various downstream tasks with substantially reduced computational cost, while qualitative results further reveal improved instance perception and completeness in visual representations. Code is available at github.com/feufhd/EinsPT.
科研通智能强力驱动
Strongly Powered by AbleSci AI