计算机科学
人工智能
Boosting(机器学习)
光学(聚焦)
代表(政治)
GSM演进的增强数据速率
匹配(统计)
感觉线索
相互信息
模式
计算机视觉
编码(集合论)
可视化
模式识别(心理学)
隐藏字幕
可视对象
结构线形
地点
模态(人机交互)
机器学习
自然语言处理
钥匙(锁)
源代码
特征学习
信息丢失
作者
Zanxi Ruan,Gao, Songqun,Kong, Qiuyu,Yiming Wang,Marco Cristani
标识
DOI:10.48550/arxiv.2602.20089
摘要
Edge-based representations are fundamental cues for visual understanding, a principle rooted in early vision research and still central today. We extend this principle to vision-language alignment, showing that isolating and aligning structural cues across modalities can greatly benefit fine-tuning on long, detail-rich captions, with a specific focus on improving cross-modal retrieval. We introduce StructXLIP, a fine-tuning alignment paradigm that extracts edge maps (e.g., Canny), treating them as proxies for the visual structure of an image, and filters the corresponding captions to emphasize structural cues, making them "structure-centric". Fine-tuning augments the standard alignment loss with three structure-centric losses: (i) aligning edge maps with structural text, (ii) matching local edge regions to textual chunks, and (iii) connecting edge maps to color images to prevent representation drift. From a theoretical standpoint, while standard CLIP maximizes the mutual information between visual and textual embeddings, StructXLIP additionally maximizes the mutual information between multimodal structural representations. This auxiliary optimization is intrinsically harder, guiding the model toward more robust and semantically stable minima, enhancing vision-language alignment. Beyond outperforming current competitors on cross-modal retrieval in both general and specialized domains, our method serves as a general boosting recipe that can be integrated into future approaches in a plug-and-play manner. Code and pretrained models are publicly available at: https://github.com/intelligolabs/StructXLIP.
科研通智能强力驱动
Strongly Powered by AbleSci AI