强化学习
对抗制
恶意软件
计算机科学
逃避(道德)
字节
黑匣子
计算机安全
人工智能
好奇心
黑客
机器学习
操作系统
社会心理学
生物
免疫学
心理学
免疫系统
作者
Dazhi Zhan,Yanyan Zhang,Ling Zhu,Jun Chen,Shiming Xia,Shize Guo,Zhisong Pan
标识
DOI:10.1016/j.aej.2024.04.024
摘要
The anti-detection capabilities of adversarial malware examples have drawn the attention of antivirus vendors and researchers. In black-box scenarios where internal information of the target model cannot be accessed, with the ability of the reinforcement learning (RL) agents to adjust strategies based on the feedback from the environment, existing RL-based methods enable evasion of Windows PE malware detectors. However, obtaining evasion rewards as positive feedback in the black-box setting proves challenging, resulting in low training efficiency. To address this issue, we introduce the intrinsic curiosity reward into the framework to motivate the agent to explore unknown state spaces and learn effective evasion strategies. Additionally, we employ the generative adversarial network (GAN) to obtain varying synthetic data, which replaces random or benign bytes as adversarial payloads for the agent's action content, improving attack capabilities and reducing the risk of hard-coded adversarial perturbations being anchored. We compare the attack performance with other RL-based baseline methods, and experimental results show that our framework is more flexible and effective, achieving a 63%-85% attack success rate against EMBER, FireEye and MalConv. Even if defensive measures are taken, the proposed method still has a certain attack capability, and the success rate remains between 48%-67%.
科研通智能强力驱动
Strongly Powered by AbleSci AI