对抗制
计算机科学
对手
公平性度量
稳健性(进化)
歪斜
结果(博弈论)
协方差
对抗性机器学习
边界判定
情感(语言学)
人工智能
机器学习
计算机安全
心理学
微观经济学
数学
统计
经济
支持向量机
电信
生物化学
化学
沟通
吞吐量
无线
基因
作者
Ninareh Mehrabi,Muhammad Naveed,Fred Morstatter,Aram Galstyan
标识
DOI:10.1609/aaai.v35i10.17080
摘要
Algorithmic fairness has attracted significant attention in recent years, with many quantitative measures suggested for characterizing the fairness of different machine learning algorithms. Despite this interest, the robustness of those fairness measures with respect to an intentional adversarial attack has not been properly addressed. Indeed, most adversarial machine learning has focused on the impact of malicious attacks on the accuracy of the system, without any regard to the system's fairness. We propose new types of data poisoning attacks where an adversary intentionally targets the fairness of a system. Specifically, we propose two families of attacks that target fairness measures. In the anchoring attack, we skew the decision boundary by placing poisoned points near specific target points to bias the outcome. In the influence attack on fairness, we aim to maximize the covariance between the sensitive attributes and the decision outcome and affect the fairness of the model. We conduct extensive experiments that indicate the effectiveness of our proposed attacks.
科研通智能强力驱动
Strongly Powered by AbleSci AI