计算机科学
水准点(测量)
解码方法
序列(生物学)
背景(考古学)
编码器
集合(抽象数据类型)
对象(语法)
匹配(统计)
人工智能
模式识别(心理学)
算法
数学
地理
遗传学
古生物学
生物
大地测量学
程序设计语言
操作系统
统计
作者
Chang Liu,Bin Zhang,Chunjuan Bo,Dong Wang
出处
期刊:Sensors
[Multidisciplinary Digital Publishing Institute]
日期:2024-07-24
卷期号:24 (15): 4802-4802
被引量:2
摘要
Query decoders have been shown to achieve good performance in object detection. However, they suffer from insufficient object tracking performance. Sequence-to-sequence learning in this context has recently been explored, with the idea of describing a target as a sequence of discrete tokens. In this study, we experimentally determine that, with appropriate representation, a parallel approach for predicting a target coordinate sequence with a query decoder can achieve good performance and speed. We propose a concise query-based tracking framework for predicting a target coordinate sequence in a parallel manner, named QPSTrack. A set of queries are designed to be responsible for different coordinates of the tracked target. All the queries jointly represent a target rather than a traditional one-to-one matching pattern between the query and target. Moreover, we adopt an adaptive decoding scheme including a one-layer adaptive decoder and learnable adaptive inputs for the decoder. This decoding scheme assists the queries in decoding the template-guided search features better. Furthermore, we explore the use of the plain ViT-Base, ViT-Large, and lightweight hierarchical LeViT architectures as the encoder backbone, providing a family of three variants in total. All the trackers are found to obtain a good trade-off between speed and performance; for instance, our tracker QPSTrack-B256 with the ViT-Base encoder achieves a 69.1% AUC on the LaSOT benchmark at 104.8 FPS.
科研通智能强力驱动
Strongly Powered by AbleSci AI