What does this writing signal mean?

Anthropic published The Capacity For Moral Self Correction In Large Language Models. This talking signal gives public context for research themes, product direction, policy, or launch framing. High-signal details: The Capacity for Moral Self-Correction in Large Language Models \ Anthropic Societal Impacts The Capacity for Moral Self-Correction in Large Language Models Feb 15, 2023.... onlylabs links this event to 1 captured evidence page and 6 related writing signals.

Anthropic Writing: The Capacity For Moral Self Correction In Large Language Models

Captured source

source ↗

anthropic.com/anthropic.com/research/the-capacity-for-moral-self-correction-in-large-language-models

The Capacity For Moral Self Correction In Large Language Models

Source ↗

published Feb 15, 2023seen 2dcaptured 8hhttp 200method plain

The Capacity for Moral Self-Correction in Large Language Models \ Anthropic Societal Impacts The Capacity for Moral Self-Correction in Large Language Models Feb 15, 2023 Read Paper

Abstract We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmful outputs -- if instructed to do so. We find strong evidence in support of this hypothesis across three different experiments, each of which reveal different facets of moral self-correction. We find that the capability for moral self-correction emerges at 22B model parameters, and typically improves with increasing model size and RLHF training. We believe that at this level of scale, language models obtain two capabilities that they can use for moral self-correction: (1) they can follow instructions and (2) they can learn complex normative concepts of harm like stereotyping, bias, and discrimination. As such, they can follow instructions to avoid certain kinds of morally harmful outputs. We believe our results are cause for cautious optimism regarding the ability to train language models to abide by ethical principles. Policy Memo Moral Self-Correction Policy Memo