计算机科学
推荐系统
联营
人工智能
代表(政治)
突出
特征学习
情报检索
图形
机器学习
协同过滤
素描
结合属性
深度学习
偏爱
卷积(计算机科学)
人机交互
外部数据表示
语义学(计算机科学)
自然语言处理
卷积神经网络
个性化
数据建模
双线性插值
任务分析
矩阵分解
语义相似性
作者
Jie Guo,Longyu Wen,Yunfei Zhao,Bin Song,Yuhao Chi
标识
DOI:10.1109/tmm.2025.3623549
摘要
Multimodal recommender systems try to integrate multimedia data (images, texts, etc.) with user-item historical records to better model user preference. However, most previous methods largely ignored the underlying fine-grained attribute features of items, which makes it difficult to fully explore users' nuanced attention across individual and combined attributes, resulting in low recommendation performance. To address these issues, this paper proposes a novel and effective self-harmonized representation learning network for multimodal recommendation, named LETTER. LETTER has the ability to effectively optimize the user and item representations for multimodal recommendation. Specifically, we design a factorized attribute interaction module that captures diverse combinations of item latent attributes using a bilinear pooling strategy. Then a dual graph convolution module is established to learn the modality-specific representations from user-item interactive and item semantic relations. Finally, we design a preference self-harmonization module that adaptively identifies the salient influencing factors of user preference, thus refining user and item representations to improve recommendation accuracy. We conduct extensive experiments on three real-world datasets, demonstrating that LETTER outperforms state-of-the-art multimodal recommendation methods.
科研通智能强力驱动
Strongly Powered by AbleSci AI