计算机科学
强化学习
Lift(数据挖掘)
背景(考古学)
推荐系统
可扩展性
软件部署
语义学(计算机科学)
人工智能
动作(物理)
链接数据
空格(标点符号)
网络拓扑
状态空间
语义网
用户建模
生产(经济)
机器学习
面子(社会学概念)
国家(计算机科学)
知识图
大数据
作者
Wang, Minmao,Xingchen Liu,Shijie Yi,L. J. Wu,Hongke Zhao,Fei Pan,Qingpeng Cai,Peng Jiang
标识
DOI:10.48550/arxiv.2510.09167
摘要
Recommender Systems (RS) are fundamental to modern online services. While most existing approaches optimize for short-term engagement, recent work has begun to explore reinforcement learning (RL) to model long-term user value. However, these efforts face significant challenges due to the vast, dynamic action spaces inherent in RS, which hinder stable policy learning. To resolve this bottleneck, we introduce Hierarchical Semantic RL (HSRL), which reframes RL-based recommendation over a fixed Semantic Action Space (SAS). HSRL encodes items as Semantic IDs (SIDs) for policy learning, and maps SIDs back to their original items via a fixed lookup during execution. To align decision-making with SID generation, the Hierarchical Policy Network (HPN) operates in a coarse-to-fine manner, employing hierarchical residual state modeling to refine each level's context from the previous level's residual, thereby reducing representation-decision mismatch. In parallel, a Multi-level Critic (MLC) provides token-level value estimates, enabling fine-grained credit assignment. Across public benchmarks and a large-scale production dataset from a leading short-video advertising platform, HSRL consistently surpasses state-of-the-art baselines. In online deployment over a 7-day A/B testing, it delivers an 18.421% ADVV lift and a 1.251% increase in Revenue, supporting HSRL as a scalable paradigm for RL-based recommendation.
科研通智能强力驱动
Strongly Powered by AbleSci AI