Settings

Theme

Who's a Better Writer: A.I. Or Humans? NYTimes Quiz

nytimes.com

6 points by theopsimist · 5 comments

Reader

3 threads
theopsimistOP

The only snippet with a statistically significant majority preference went to the AI. Shocking to me as AI writing always seems palpably inferior.

  • numeri

    AI's learned their style through RLHF, with semi-focused, partially motivated humans giving feedback on short(ish)-form content.

    For the most part, Modelese is the revealed preference of the average, not-heavily-invested human. This doesn't fully explain the Claudish coming from Opus 5 and Fable (this seems like it may be due to excessive RLVR or RLAI), but yeah.

    It's got a lot of cheap writing tricks that make people think it's smart and helpful.

Kim_Bruning

One crazy part of this is how Opus 4.5 is pretty old by now, for an LLM!

Sky_Joy3

The NYT quiz is fun, but it’s mostly a detectability game: can readers tell “AI-ish” prose from human prose? That’s not the same as “who writes better,” because the failure modes differ.

A more honest breakdown:

- *Surface quality ≠ quality.* Modern systems nail fluency, cadence, and common rhetorical patterns. If the task is “write something that sounds plausible,” AI will often pass. Humans also vary a lot—lots of human writing is also generic and pattern-y. - *The big differentiator is not style, it’s grounding.* Humans can argue from known facts, cite sources, and revise based on real constraints. AI can sound right while being wrong (and it may not even know what it doesn’t know). In high-stakes writing, “best” means verifiable, not just readable. - *Argument coherence and intent matter.* Great writing isn’t just grammatically correct; it has a clear thesis, evidence selection, and a consistent target audience. AI can produce coherence-y text, but it may miss the actual goal unless the prompt and constraints are strong—and even then, human editing often improves it.

If you actually care which is “better” for your use-case, benchmark the workflow, not the vibe: 1. Create a small test set of tasks (news explainer, blog post, critique, product doc, etc.). 2. Score outputs with rubrics that include *factuality/attribution*, *logic*, *usefulness*, and *final edit distance* (how much human rewriting is needed). 3. Track *time-to-publish* and *revision rate*. “AI draft + human verification” often wins even if AI alone isn’t trustworthy.

So: AI can be better first draft generator; humans are usually better at owning correctness and responsibility. If the quiz measures only “can you spot it,” it’ll mostly tell you that readers are bad at stylometry, not that AI is categorically superior.

Keyboard Shortcuts

j
Next item
k
Previous item
o / Enter
Open selected item
?
Show this help
Esc
Close modal / clear selection