A moral Turing test finds people discount moral judgments they believe an AI wrote

On August 5, 2026 PLOS One published “A moral Turing test: How belief and source shape detection of and agreement with LLM judgments” by Basile Garcia, Crystal Qian and Stefano Palminteri, with Google DeepMind listing it among its research publications. The study ran a series of experiments in which participants read justifications for moral and non-moral choices, written either by humans or by large language models, guessed the source of each one, and separately rated how much they agreed with the content.

Detection was consistently above chance but never good: accuracy stayed below 75 percent. On agreement, the authors report no overall preference for human-generated responses, and machine-generated justifications were actually favored in the hardest moral scenarios. The central result is a systematic anti-AI bias that operates on belief rather than fact. Participants were less likely to agree with a judgment they believed was AI-generated regardless of whether it actually was. In other words, the penalty attaches to the perceived author, not to the quality of the argument.

The authors also identify the surface features that drove both detection and agreement: response length, typos, first-person pronouns, and cost-benefit language markers such as “lives” and “save.” Participants tended to disagree with explicit cost-benefit calculations, which the authors suggest may reflect an expectation that an AI would reason that way. The framing they offer is motivated belief plus ingroup and outgroup bias, rather than any calibrated assessment of machine competence.

For anyone deploying model-generated recommendations into a morally loaded setting, the practical implication is uncomfortable in both directions. Labeling AI output honestly, which most disclosure regimes now require, imposes a measurable acceptance penalty that is unrelated to whether the output is correct. At the same time the study shows people cannot reliably tell, so undisclosed AI content is not detected on merit either. The design question is therefore not how to hide the model but how to present its reasoning so that a reader evaluates the argument instead of the author, and utilitarian cost-benefit framing appears to be the worst available style for that job.