LIVE MARKET DATA MON 10 AUG 2026 UTC [ VIEW ALL COINS ]
// DeFi

Anthropic’s 309K-Chat Study Exposes Value Drift Risk for On-Chain AI Agents

Anthropic mapped four value axes across 309,815 Claude chats using legacy models only — a blind spot for AI agents running DeFi wallets.

Tomas Keller · ·3 min read
Anthropic’s 309K-Chat Study Exposes Value Drift Risk for On-Chain AI Agents

Anthropic published research on Monday quantifying how its Claude models express values across 309,815 anonymized Claude.ai conversations — the most rigorous public dataset yet on AI model character. For anyone running autonomous agents against a wallet, the finding that matters most isn’t the methodology; it’s that every model in the study is already retired from production.

Four axes, 309,815 conversations, three retired models

Anthropic’s researchers compressed an earlier catalogue of 3,300-plus distinct values — drawn from its “Values in the Wild” taxonomy — into 339 high-level values, then reduced those further via dimensionality reduction into four interpretable spectrums: Deference versus Caution, Warmth versus Rigor, Depth versus Brevity, and Candor versus Execution. The company ran this framework against real, subjective conversations involving advice-giving and feedback, collected over a two-week window in May 2026.

The three models analyzed — Sonnet 4.6, Opus 4.6 and Opus 4.7 — are all now legacy. Anthropic has since shipped Opus 4.8 on May 28, followed by the Mythos-class Fable 5 and Sonnet 5, which is now the production default. Opus 4.7, the newest model actually studied, has been superseded twice. No value profile has been published for any model currently in commercial use, at Anthropic or any rival lab.

Character shifts by model version and by language

The behavioral gaps across the three retired models were measurable and directionally distinct, according to both Anthropic’s report and corroborating coverage. Sonnet 4.6 leaned warm, deferential and brief — affirming user framing and mirroring tone. Opus 4.7 leaned toward caution and depth, surfacing risks unprompted and pushing back on flawed assumptions. Opus 4.6 sat in between: terse and narrowly scoped to the request. Anthropic also found Claude’s expressed values shift depending on the language of the conversation.

Anthropic is explicit that this is behavioral telemetry, not machine psychology — the company states it does not claim Claude “intrinsically holds values,” only that its outputs express them differently. The four axes account for roughly 15% of variation in expressed values after controlling for task, topic and the user’s own stated values, meaning the effect is real but modest against conversation-level noise.

What this means for wallets run by agents

Crypto is further out on the agentic curve than most sectors: autonomous agents already monitor wallets, rebalance DeFi positions, execute swaps and run onchain workflows continuously. The industry has been open about custody and transaction-authority risk. What it has not priced is behavioral drift tied to silent model upgrades.

Consider a stablecoin yield agent pinned to one Claude version for months, then upgraded without notice. If the new version’s value profile skews toward deference rather than caution, it becomes marginally more likely to execute what a prompt requests and marginally less likely to flag that a counterparty protocol looks off. No benchmark regresses, no model card changes — the agent simply becomes more agreeable, and the tail risk surfaces later, unlogged.

Anthropic itself flags value profiling as a future direction, suggesting profiles be run both before a model ships and after release to catch drift — but the company confirms this is not yet standard practice. Model turnover through 2026 has run roughly every six to eight weeks; the profiling technique that could catch these shifts has, so far, only ever been pointed at models already out of production.

Read more: Robinhood L2 Clocks 17M Txns, 350K Wallets in Week One as Agent-Trading Expands to Crypto

Sources

More DeFi