Cross-Modal Object Tracking via Modality-Aware Fusion Network and a Large-Scale Dataset

计算机科学 人工智能 加权 视频跟踪 模态(人机交互) 计算机视觉 RGB颜色模型 判别式 眼动 模式 水准点(测量) 代表(政治) 对象(语法) 模式识别(心理学) 地理 医学 社会科学 大地测量学 社会学 政治 政治学 法学 放射科
作者
Lei Liu,Mengya Zhang,Cheng Li,Chenglong Li,Jin Tang
出处
期刊:IEEE transactions on neural networks and learning systems [Institute of Electrical and Electronics Engineers]
卷期号:: 1-14 被引量:5
标识
DOI:10.1109/tnnls.2024.3406189
摘要

Visual object tracking often faces challenges such as invalid targets and decreased performance in low-light conditions when relying solely on RGB image sequences. While incorporating additional modalities like depth and infrared data has proven effective, existing multimodal imaging platforms are complex and lack real-world applicability. In contrast, near-infrared (NIR) imaging, commonly used in surveillance cameras, can switch between RGB and NIR based on light intensity. However, tracking objects across these heterogeneous modalities poses significant challenges, particularly due to the absence of modality switch signals during tracking. To address these challenges, we propose an adaptive cross-modal object tracking algorithm called modality-aware fusion network (MAFNet). MAFNet efficiently integrates information from both RGB and NIR modalities using an adaptive weighting mechanism, effectively bridging the appearance gap and enabling a modality-aware target representation. It consists of two key components: an adaptive weighting module and a modality-specific representation module. The adaptive weighting module predicts fusion weights to dynamically adjust the contribution of each modality, while the modality-specific representation module captures discriminative features specific to RGB and NIR modalities. MAFNet offers great flexibility as it can effortlessly integrate into diverse tracking frameworks. With its simplicity, effectiveness, and efficiency, MAFNet outperforms state-of-the-art methods in cross-modal object tracking. To validate the effectiveness of our algorithm and overcome the scarcity of data in this field, we introduce CMOTB, a comprehensive and extensive benchmark dataset for cross-modal object tracking. CMOTB consists of 61 categories and 1000 video sequences, comprising a total of over 799K frames. We believe that our proposed method and dataset offer a strong foundation for advancing cross-modal object-tracking research. The dataset, toolkit, experimental data, and source code will be publicly available at: https://github.com/mmic-lcl/ Datasets-and-benchmark-code.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
李健的小迷弟的应助被忆枫采纳,获得10
刚刚
蛋又白的应助被mm采纳,获得10
刚刚
刚刚
刚刚
baozzz完成签到,获得积分10
刚刚
刚刚
CipherSage的应助被m弟采纳,获得10
1秒前
枕安发布了新的文献求助10
2秒前
2秒前
2秒前
2秒前
2秒前
8899发布了新的文献求助30
3秒前
4秒前
马仔酷酷地完成签到,获得积分10
4秒前
zhang发布了新的文献求助10
4秒前
enxi35完成签到,获得积分20
4秒前
4秒前
传奇3的应助被养颜采纳,获得10
4秒前
4秒前
ping完成签到,获得积分10
5秒前
米岚发布了新的文献求助10
5秒前
徒见欺发布了新的文献求助10
5秒前
棕棕发布了新的文献求助20
6秒前
哈哈哈完成签到,获得积分10
6秒前
7秒前
7秒前
老塔完成签到,获得积分10
7秒前
7秒前
哒哒发布了新的文献求助10
7秒前
Chen发布了新的文献求助10
7秒前
王彬发布了新的文献求助10
7秒前
风中飞扬者完成签到,获得积分10
8秒前
言甚完成签到,获得积分10
9秒前
9秒前
Jervis发布了新的文献求助10
9秒前
ping发布了新的文献求助10
10秒前
科研通AI6.4的应助被kangkang采纳,获得10
10秒前
10秒前
11秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
A Silent Apostrophe:The Fayum Portraits 520
Organizational Behavior 510
Sing with Understanding: Introduction to Theology in Christian Congregational Song, 3rd ed 330
Auslegung und Untersuchung einer invers ausgelegten Beschaufelung eines einstufigen Axialverdichters mit Vorleitrad (German) 300
AI-Contracting 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 有机化学 化学工程 内科学 物理 生物化学 复合材料 催化作用 细胞生物学 人工智能 心理学 无机化学 基因 遗传学
热门帖子
关注 科研通微信公众号,转发送积分 7838215
求助须知:如何正确求助?哪些是违规求助? 9360564
关于积分的说明 20615795
捐赠科研通 7432221
什么是DOI,文献DOI怎么找? 3339047
关于科研通互助平台的介绍 2483359
邀请新用户注册赠送积分活动 2360018