Project Swap: 201 Anthropic employees let Claude agents trade books for them

On September 24, 2026, Anthropic published Project Swap, a controlled experiment in agent-to-agent commerce. Over the summer, 201 Anthropic employees in six offices (San Francisco, New York, London, Seattle, Washington DC and Dublin) each brought a book to give away and spent about five minutes telling a Claude agent what they wanted to read. Claude (Fable 5 for this step) turned each conversation into a ranking of every book in that person’s local pool, and the agents then bartered on a decentralized trading floor with time and message limits. Participants separately ranked 10 books themselves to provide ground truth. Anthropic then re-ran the market dozens of times, varying the model (Haiku 4.5, Sonnet 4.5, Opus 4.8, Fable 5) and the instructions (“ruthless” versus “prosocial”).

From a five-minute chat, Claude’s ranking agreed with a person’s own on 61 percent of book pairs, against 50 percent for chance. People ended up at 0.55 on their own normalized ranking scale, roughly their fifth choice out of ten, while the best possible assignment would have scored 0.89. Decomposing the gap, the best assignment computed from Claude’s rankings but scored on people’s real preferences reached only 0.60, so imprecise preference capture accounted for about 85 percent of the shortfall and the free-for-all market design for the rest. Model strength mattered more than instructions: measured on Claude’s rankings, Haiku trading floors averaged 0.75 and Opus floors 0.88, while ruthless agents beat prosocial ones only slightly. Among those who responded, average satisfaction with the book received was 7.2 out of 10, and participants said they would let an agent spend about 30 percent of a yearly book budget unsupervised.

The authors list the limits themselves. The sample is Anthropic employees, who are probably more willing than most people to trust Claude; there were no incentives to rank carefully; every agent was cooperative, with no adversarial or deceptive counterparties; the market rules were fixed; some books were never delivered; and satisfaction data comes from the subset who answered the endline survey.

The finding that matters for anyone building delegated agents is where the value leaked. Negotiation was the easy part; understanding what the principal actually wants was the bottleneck, and people who wrote more in intake were better represented. What the study does not show is how agents behave in high-stakes markets, against adversaries, or on behalf of people who do not already trust the model.