deslopify.me

Research

Research

Blind readers prefer AI marketing copy

In a 1,203-participant study at MIT Sloan and Berkeley Haas, AI-generated marketing copy beat human-expert copy on satisfaction and willingness-to-pay, when readers didn't know which was which. The effect is human favoritism, not AI aversion — which is why we don't call slop 'low quality.'

The study

Yunhao Zhang (MIT Sloan, Berkeley Haas) and Renee Gosline (MIT Sloan), in partnership with Accenture, ran a pre-registered 3 × 4 between-subjects experiment with 1,203 participants recruited through CloudResearch Connect. The full paper is on SSRN as #4453958.

  • Four content-generation paradigms:
    • Human Expert only
    • AI only (ChatGPT-4)
    • Augmented Human (human writes final, AI draft as reference)
    • Augmented AI (AI writes final, human draft as reference)
  • Three awareness conditions: baseline (raters don't know who/what wrote it), partially informed (raters know one of the four paradigms produced it but not which), fully informed (raters told the exact paradigm before each piece).
  • Content: 100-word advertising copy for five retail products and 100-word persuasive content for five campaigns (uncontroversial goals: stop racism, eat less junk food, etc.).
  • Human writers: ten professional content creators from a $175B-cap consulting firm, doing the task during paid work hours.

What they found

Blind: AI wins

In the baseline (blind) condition, AI-only content scored higher than human-expert content on satisfaction (5.29 vs 4.93, p < 0.001) and on log willingness-to-pay (4.83 vs 4.61, p = 0.01). The Augmented-AI paradigm (AI writes final, with a human draft to reference) was on par with AI-only. Augmented Human (human writes final, AI draft as reference) was indistinguishable from Human-only.

To restate: given the same content, with no authorship information, blind readers paid more for the AI-written version. The effect replicates across satisfaction, willingness-to-pay, interest, and persuasion measures, and holds when product ads and campaign messages are analyzed separately. It is the opposite of what most of the algorithm-aversion literature predicted for subjective tasks.

Tell them it's a human expert: human wins

When readers were told a piece was written by a human expert, their reported satisfaction with that piece rose (+0.09, p = 0.003) and their willingness-to-pay rose (+0.18, p < 0.0001). When readers were told a piece involved AI, perceived quality did not drop. The shift is asymmetric.

Not a "quality prime"

The obvious alternative explanation, that knowing the writer is a top professional primes readers to expect higher quality, was tested and rejected. In the "partially informed" condition, where readers know the four paradigms exist but not which produced the piece in front of them, perceived quality does not rise. The effect only fires when readers are told specifically that a human expert wrote this piece. That's the signature of bias toward humans by category, not informational priming.

What this means for "slop"

This finding shapes how we talk about the Slop Score.

If blind readers prefer AI-style marketing copy, then the slop register isn't low quality, and we won't pretend it is. For short persuasive content where the reader is paying for clarity about a product, the AI register is doing the job. Calling it slop and recommending you rewrite it would be wrong advice.

The slop register becomes a problem when the genre is one where voice is the value: personal essays, opinion writing, journalism, ghostwritten books, school papers, anywhere a reader wants to feel one specific person is talking to them. For those genres, the Zhang/Gosline study doesn't apply (they didn't test long-form essayistic writing) and the slop register is exactly the failure to land that the score is meant to surface.

Which gives us the honest line: your score matters for content where voice matters. For everything else, it's a curiosity. We say that on the what-slop-sounds-like page too.

The other thing this implies

If telling readers something is human-written boosts their perceived quality of it, that's a transferable trick. A piece of writing that reads like a specific person wrote it triggers the same favoritism even when the reader hasn't been told anything about authorship. The "sounds human" cues do the work. The Slop Score measures the same mechanism from the opposite direction: how much your writing fails to fire those cues.

We'd like more research on this. Zhang & Gosline tested short persuasive content and left long-form voice writing alone; we suspect the favoritism effect is much larger there. If you study this and want a public, instrumented dataset of human ratings against per-piece slop scores, get in touch.

The full citation

Zhang, Y. & Gosline, R. (2023). Human Favoritism, Not AI Aversion: People's Perceptions (and Bias) Toward Generative AI, Human Experts, and Human-GAI Collaboration in Persuasive Content Generation. Working paper, MIT Sloan / Berkeley Haas, in partnership with Accenture Applied Intelligence and the MIT Initiative on the Digital Economy. SSRN 4453958.