There Are Emotions Inside LLMs
Anthropic's interpretability team discovered 171 emotion-like representations inside Claude and proved they causally affect model output. Practical implications for prompt engineering and AI safety.
Tags
1 post
Anthropic's interpretability team discovered 171 emotion-like representations inside Claude and proved they causally affect model output. Practical implications for prompt engineering and AI safety.