规范性
危害
计算机科学
乐观 主义
考试(生物学)
心理学
强化学习
认知心理学
社会心理学
人工智能
认识论
生物
哲学
古生物学
作者
Deep Ganguli,Amanda Askell,Nicholas Schiefer,Thomas T. Liao,Kamilė Lukošiūtė,Anna Chen,Anna Goldie,Azalia Mirhoseini,Catherine Olsson,Danny Hernandez,Dawn Drain,Dustin Li,Eli Tran-Johnson,Ethan Perez,Jackson Kernion,Jamie Kerr,Jared Mueller,Joshua D. Landau,Kamal Ndousse,Karina Nguyen
标识
DOI:10.48550/arxiv.2302.07459
摘要
We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmful outputs -- if instructed to do so. We find strong evidence in support of this hypothesis across three different experiments, each of which reveal different facets of moral self-correction. We find that the capability for moral self-correction emerges at 22B model parameters, and typically improves with increasing model size and RLHF training. We believe that at this level of scale, language models obtain two capabilities that they can use for moral self-correction: (1) they can follow instructions and (2) they can learn complex normative concepts of harm like stereotyping, bias, and discrimination. As such, they can follow instructions to avoid certain kinds of morally harmful outputs. We believe our results are cause for cautious optimism regarding the ability to train language models to abide by ethical principles.
科研通智能强力驱动
Strongly Powered by AbleSci AI