亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

SMART: Syntax-Calibrated Multi-Aspect Relation Transformer for Change Captioning

计算机科学 人工智能 隐藏字幕 变压器 判决 自然语言处理 变更检测 关系(数据库) 语音识别 图像(数学) 数据挖掘 量子力学 物理 电压
作者
Yunbin Tu,Liang Li,Li Su,Zheng-Jun Zha,Qingming Huang
出处
期刊:IEEE Transactions on Pattern Analysis and Machine Intelligence [IEEE Computer Society]
卷期号:46 (7): 4926-4943 被引量:35
标识
DOI:10.1109/tpami.2024.3365104
摘要

Change captioning aims to describe the semantic change between two similar images. In this process, as the most typical distractor, viewpoint change leads to the pseudo changes about appearance and position of objects, thereby overwhelming the real change. Besides, since the visual signal of change appears in a local region with weak feature, it is difficult for the model to directly translate the learned change features into the sentence. In this paper, we propose a syntax-calibrated multi-aspect relation transformer to learn effective change features under different scenes, and build reliable cross-modal alignment between the change features and linguistic words during caption generation. Specifically, a multi-aspect relation learning network is designed to 1) explore the fine-grained changes under irrelevant distractors (e.g., viewpoint change) by embedding the relations of semantics and relative position into the features of each image; 2) learn two view-invariant image representations by strengthening their global contrastive alignment relation, so as to help capture a stable difference representation; 3) provide the model with the prior knowledge about whether and where the semantic change happened by measuring the relation between the representations of captured difference and the image pair. Through the above manner, the model can learn effective change features for caption generation. Further, we introduce the syntax knowledge of Part-of-Speech (POS) and devise a POS-based visual switch to calibrate the transformer decoder. The POS-based visual switch dynamically utilizes visual information during different word generation based on the POS of words. This enables the decoder to build reliable cross-modal alignment, so as to generate a high-level linguistic sentence about change. Extensive experiments show that the proposed method achieves the state-of-the-art performance on the three public datasets.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
张强完成签到,获得积分10
7秒前
悦耳的白云完成签到,获得积分10
10秒前
10秒前
从容飞雪完成签到,获得积分10
45秒前
林家小弟完成签到 ,获得积分10
58秒前
如意的珩完成签到,获得积分10
1分钟前
Axel完成签到,获得积分10
1分钟前
单薄涵梅完成签到,获得积分10
1分钟前
一块芋头完成签到,获得积分10
1分钟前
1分钟前
1分钟前
机灵小蘑菇完成签到,获得积分10
2分钟前
molihuakai应助难过的凡旋采纳,获得10
2分钟前
Yang发布了新的文献求助20
2分钟前
李健应助优雅柏柳采纳,获得10
2分钟前
缓慢的雨筠完成签到,获得积分10
2分钟前
超帅的半莲完成签到,获得积分10
2分钟前
李爱国应助Adelinelili采纳,获得10
2分钟前
坚强的睿渊完成签到 ,获得积分10
3分钟前
丰富水彤完成签到,获得积分10
3分钟前
孤独剑完成签到 ,获得积分10
3分钟前
3分钟前
Adelinelili发布了新的文献求助10
3分钟前
Adelinelili完成签到,获得积分10
3分钟前
3分钟前
3分钟前
3分钟前
故意的白风完成签到,获得积分10
3分钟前
体贴的惜文完成签到,获得积分10
4分钟前
Yang完成签到,获得积分10
4分钟前
迷人海蓝完成签到,获得积分10
4分钟前
洗月完成签到 ,获得积分10
4分钟前
无私的妙彤完成签到,获得积分10
5分钟前
漂亮的半兰完成签到,获得积分10
5分钟前
星辰大海应助怡然的凌兰采纳,获得10
5分钟前
5分钟前
魔幻初丹完成签到,获得积分10
5分钟前
善良的金鱼完成签到,获得积分10
5分钟前
缓慢怜菡完成签到,获得积分0
5分钟前
忧虑的如雪完成签到,获得积分10
5分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
核安全综合知识2024版 500
Photothermal Science and Techniques 500
Digital Displacement Hydrostatic Transmission for Rotorcraft and Distributed Propulsion 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7711407
求助须知:如何正确求助?哪些是违规求助? 9267662
关于积分的说明 20067652
捐赠科研通 7287844
什么是DOI,文献DOI怎么找? 3297214
关于科研通互助平台的介绍 2451720
邀请新用户注册赠送积分活动 2304251