计算机科学
推论
判别式
人工智能
稳健性(进化)
机器学习
代表(政治)
延迟(音频)
软件部署
样品(材料)
数据挖掘
还原(数学)
基线(sea)
语义学(计算机科学)
钥匙(锁)
自然语言处理
图像检索
功率(物理)
对比度(视觉)
审问
构造(python库)
标识
DOI:10.1109/icmlca66850.2025.11336577
摘要
Cross-modal retrieval (CMR) has long grappled with a challenging trade-off between performance and efficiency. Mainstream approaches typically depend on computationally intensive dual-encoder architectures, such as ResNet-50 and BERT, which lead to high inference latency and significant deployment costs. To tackle this issue, the present paper introduces a lightweight cross-modal retrieval framework grounded in Hard Negative Sample Mining (HNM). This framework utilizes MobileNetV3 and DistilBERT as image and text encoders, respectively, thereby preserving robust representational capabilities while markedly reducing computational complexity. Specifically, the HNM contrastive loss dynamically identifies and emphasizes the most hard negative samples within each batch. This mechanism allows the model to concentrate on optimizing decision boundaries within a shared semantic space, effectively compensating for the limited discriminative power of lightweight encoders. Experimental results indicate that this framework achieves approximately 46.7 % reduction in parameters and 17.9 % decrease in inference latency on the Flickr30k dataset while surpassing traditional heavyweight baseline models across multiple Recall@K metrics. This work illustrates that synergistic optimization of lightweight architectures alongside sophisticated negative sample mining techniques facilitates efficient and accurate cross-modal retrieval under resource-constrained conditions, offering an effective pathway for future research aimed at balancing performance with deployability.
科研通智能强力驱动
Strongly Powered by AbleSci AI