后门
对手
可见的
计算机科学
对抗制
动作(物理)
多智能体系统
分布式计算
利用
强化学习
软件部署
计算机安全
情感(语言学)
人工智能
作者
Shuo Chen,Yue Qiu,Jie Zhang
标识
DOI:10.1109/tdsc.2025.3620528
摘要
Backdoor attacks on reinforcement learning implant a backdoor in a victim's policy. Once the victim observes trigger signals, it switches to the abnormal mode and fails its task. Most existing attacks assume the adversary can arbitrarily modify the victim's observations (i.e., vision-perturbation), which may be impractical. A more practical way is letting one adversary agent use its actions to affect the victim's observations (i.e., action-perturbation). However, in partially observable multiagent systems, agents may not always be able to observe others. When and how much the adversary agent can affect others' observations are uncertain. Also, we want the adversary agent to trigger others' backdoors with a few actions for stealthiness. To achieve that, we first design a novel training framework to produce auxiliary rewards that measure the influence of each trigger action on others' observations. We then use these rewards to train a trigger policy that guides the adversary agent to efficiently affect others' observations. Given the affected observations, we train the other agents to perform abnormally. Experiments show that the proposed method enables the adversary agent to trigger others' backdoors efficiently. Finally, we study the defense methods against the attack and provide insights for the deployment of multiagent systems.
科研通智能强力驱动
Strongly Powered by AbleSci AI