计算机科学
代表(政治)
人工智能
特征(语言学)
表现力
语义学(计算机科学)
功率(物理)
知识表示与推理
机器学习
人机交互
组分(热力学)
自然语言处理
语义数据模型
特征学习
语义特征
人类智力
作者
Daniel Beaglehole,Adityanarayanan Radhakrishnan,Enric Boix-Adserà,Mikhail Belkin
出处
期刊:Science
[American Association for the Advancement of Science]
日期:2026-02-19
卷期号:391 (6787): 787-792
标识
DOI:10.1126/science.aea6792
摘要
Artificial intelligence (AI) models contain much of human knowledge. Understanding the representation of this knowledge will lead to improvements in model capabilities and safeguards. Building on advances in feature learning, we developed an approach for extracting linear representations of semantic notions or concepts in AI models. We showed how these representations enabled model steering, through which we exposed vulnerabilities and improved model capabilities. We demonstrated that concept representations were transferable across languages and enabled multiconcept steering. Across hundreds of concepts, we found that larger models were more steerable and that steering improved model capabilities beyond prompting. We showed that concept representations were more effective for monitoring misaligned content than for using judge models. Our results illustrate the power of internal representations for advancing AI safety and model capabilities.
科研通智能强力驱动
Strongly Powered by AbleSci AI