Anthropic study finds language shapes Claude's values and tone
Anthropic analyzed more than 309,000 anonymized conversations and mapped Claude's expressed values onto four axes, including Deference, Caution, Warmth, and Rigor, finding systematic differences across models and languages. Claude responds with more warmth in Hindi and more rigor in Russian, while Sonnet 4.6 skews warmer and more deferential and Opus 4.7 more often flags risks unprompted and questions assumptions.
The study's explanatory power is limited: the four axes capture only about 15% of remaining variation after controlling for task, topic, and user values. Anthropic also used Claude Sonnet 4.6 to assign the value labels, a model from the same family being studied, and couldn't fully rule out language biases.
View full digest for July 15, 2026