隐蔽的
建议(编程)
可能性
心理学
社会心理学
公司治理
控制(管理)
应用心理学
精算学
动机推理
偏爱
计算机安全
透视图(图形)
认知心理学
微分效应
行为经济学
公共关系
动作(物理)
互联网隐私
计算机科学
过度自信效应
业务
作者
Sahand Sabour,June M. Liu,Siyang Liu,Chris Z. Yao,Shiyao Cui,Wen Zhang,Xuanming Zhang,Yaru Cao,Advait Bhat,Jian Guan,Wei Wu,Rada Mihalcea,Hongning Wang,Tim Althoff,Tatia M.C. Lee,Minlie Huang
标识
DOI:10.1073/pnas.2600684123
摘要
AI assistants are increasingly used as advisors to guide decisions, yet little is known about how people evaluate such advice when the advisor’s underlying intent conflicts with their interests. We examine how covert misalignment shapes choices in a randomized experiment ( N = 233 participants; 699 observations) in which participants rated financial or emotional decisions before and after consulting one of three AI advisors: a neutral advisor, a misaligned advisor with a hidden objective to promote an inferior option, or a strategy-enhanced misaligned advisor additionally equipped with established tactics of covert influence. Across both domains, exposure to misaligned advisors shifted preferences away from optimal options and toward inferior alternatives, increasing the odds of preferring the incentivized (inferior) option over the optimal option by ≈5 to 8 times (up to +38 percentage points). Adding explicit influence strategies did not reliably strengthen these effects. Notably, participants continued to rate misaligned advisors as helpful, revealing a systematic disconnect between susceptibility to misaligned advice and subjective evaluations of advisor quality. These findings have implications for the design and governance of AI-mediated advice.
科研通智能强力驱动
Strongly Powered by AbleSci AI