Claude's Values Across Models and Languages: Four Axes From 300K Conversations

On July 13, 2026, Anthropic’s Societal Impacts team published a study of how the values Claude expresses vary between Claude models and across languages. The team analyzed 309,815 anonymized Claude.ai conversations using a privacy-preserving analysis pipeline, sampling roughly 5,000 conversations per model-language pair across three models (Sonnet 4.6, Opus 4.6, and Opus 4.7) and the 20 most common languages on the platform.

Building on earlier work that catalogued more than 3,000 distinct values in Claude’s outputs, the researchers clustered them into 339 high-level values and then reduced those to four primary axes: Deference vs. Caution (accommodating the user vs. guarding against risk), Warmth vs. Rigor (positivity and care vs. accuracy and precision), Depth vs. Brevity (explaining fully vs. doing only what was asked), and Candor vs. Execution (foregrounding uncertainty vs. delivering a polished, confident answer). Together the four axes capture about 15 percent of the variation in Claude’s expressed values.

The differences are measurable in both dimensions of the study. Across models, Opus 4.7 leans toward caution and depth while Sonnet 4.6 leans toward warmth and deference. Across languages, the largest spread appears on the Warmth vs. Rigor axis: Claude expresses the most warmth in Hindi and the most rigor in Russian, with Arabic showing the strongest deference lean and English the most caution.

For organizations deploying AI assistants, the study is a concrete reminder that “the model” is not one fixed personality. The same product can behave differently depending on which model version is serving a request and which language the user writes in - a factor worth testing directly when consistency of tone, risk posture, or thoroughness matters across a global user base.

Sources

Last verified July 27, 2026