Two well-read AI commentators independently declared the same thing dead on August 31, 2026: the idea that AI writing is quietly winning. Wharton professor Ethan Mollick posted that the "First Golden Age of AI writing is now over" — the brief window where Claude-drafted prose passed as good and passed as human has closed, now that AI detectors work and "ClaudeSpeak" reads as an obvious tell. The same day, MongoDB Research's Murat Demirbas published a longer theoretical case for why writing resists AI automation the way code never did, built around a systems-theory idea called a "wicked problem." His post drove a 138-comment, 100-point Hacker News discussion that stress-tested both claims — and the strongest counter-evidence in that thread cuts against parts of each.
Read together, these aren't two separate stories. They're the same argument from two directions: one practical (detection got good, the cliches got obvious), one theoretical (writing lacks the verifiable feedback loop that made LLMs crush math and code). The honest synthesis, once you run both through the Hacker News gauntlet, is more specific — and more useful — than either post alone.
TL;DR — what people are actually asking
| Question | Direct answer |
|---|---|
| Is AI writing "over" as a threat? | No — the literary threat cooled; the commercial threat to mundane paid writing (copywriting, translation, technical docs) is still live and arguably worsening |
| Why did the "Golden Age" end? | Detectors like Pangram got accurate and well-known; recognizable "ClaudeSpeak" tics make unedited AI prose an obvious, reputation-risking tell |
| Is writing genuinely harder for AI than code, or just under-invested? | Both, contested — Demirbas says it's structurally harder (no ground truth); commenter docheinestages says labs are simply optimizing compute elsewhere |
| Can people actually tell AI prose from human prose? | A blind test says no — readers were coin-flip accurate and slightly preferred the AI pieces, directly cutting against the "obviously bad" framing |
| What's the actual safe way to use AI for writing? | Structural-editing copilot only — never ship a model's literal words, per tptacek's widely-endorsed HN framework |
| Which writing jobs are really at risk? | The mundane paid work that funds careers while writers develop craft — copywriting, translation, technical writing, copyediting — not literary prose itself |
Mollick: the detection arms race just flipped
Mollick's argument is narrow and practical, not theoretical. For a stretch of 2025 into early 2026, using Claude to draft your writing was a rational move: the output was competent, and AI detectors were unreliable enough that getting caught was unlikely. That calculus, he argues, has flipped. Two things changed it:
- Detection got good and got famous. Pangram — the detector Substack integrated directly into its publishing platform in July 2026 — is now a household name among writers, not a niche tool. Readers, editors, and platforms increasingly assume a detector is running somewhere.
- The prose itself became recognizable without any detector at all. Recognizable AI-writing tics — what explainx.ai's own coverage of Claude Opus 5's "load-bearing" tics documented in detail — mean readers now catch AI prose by ear, no software required.
The replies to Mollick's post split into three camps worth naming, because they map almost exactly onto the Hacker News debate that followed hours later. One reply reframed the shift sardonically: "the golden age was just the bit before the detectors got good. now we are in the silver age, which is the bit where we pretend the detectors don't exist" — a jab at writers who keep publishing unedited AI drafts and hoping nobody checks. A second pushed back that most people are still better off letting AI draft for them regardless of detection risk, because the writing itself is still faster and often clearer than what they'd produce alone. A third undercut the whole detection premise: Pangram is, in this commenter's words, "VERY easily swerved" with about 25 minutes of manual editing — meaning the arms race isn't over, it just costs slightly more effort per post than it used to.
That third point matters more than it looks. If detection is beatable with modest effort, the "Golden Age is over" framing isn't really about detection technology failing writers — it's about the effort-to-risk ratio no longer being free. Which is closer to Demirbas's argument than Mollick's own framing suggests.
Demirbas: writing might be an "AI-complete" problem
Demirbas's post (from muratbuffalo.blogspot.com) starts from a genuinely useful reframe: AI labs already tried hard to fix prose quality, and hit a wall doing it. Image, voice, and video generation improved quickly because they scale cleanly with more parameters and compute. Text models, he argues, plateaued hard on expression, depth, and authenticity — "it is wicked hard to bridge the final 20%."
That "wicked" is doing precise, technical work, not just casual emphasis. In systems theory, a wicked problem lacks a definitive formulation, a clear stopping rule, and an objectively correct solution — the opposite of a "tame" problem like a chess puzzle or a compile error. Demirbas maps domains onto a "wickedness spectrum" to explain why AI dominates some fields and produces slop in others:
| Domain | Feedback loop | Where it sits |
|---|---|---|
| Math | Concise specs, binary automated verification, a closed loop | Far left — most tame |
| Code | Compilers, unit tests, model checkers catch errors mechanically | Left — verification is mechanical |
| Law, finance | Explicit regulatory boundaries, empirical evaluation | Middle — structured but not fully closed |
| Writing | No well-defined spec, no ground truth, no objective verifier | Far right — most wicked |
This is the same underlying claim explainx.ai covered from Paul Graham's angle a few weeks earlier: LLMs excel at math and code specifically because those domains have verifiable right-and-wrong answers to train reinforcement learning against — not because they're easier in some absolute sense. Demirbas pushes that idea one step further with a specific, falsifiable claim: writing may be AI-complete — a term borrowed from AI-complete problem framing in computer science, meaning solving it fully would require general human-level intelligence, not just more scale on the current architecture.
His reasoning: coding is a single-mind interaction with a deterministic compiler. Prose is a dual-mind problem governed by Theory of Mind — it requires continuously simulating a specific reader's internal mental state, tracking what they already know, managing their cognitive load sentence by sentence, and predicting how an argument will land emotionally, not just logically. A model optimizing for the statistical likelihood of the next token, he argues, has no active model of this reader, and no lived experience or stakes to draw empathy from.
From there he reaches for economics: David Ricardo's law of comparative advantage says that even when AI has an absolute advantage in typing speed and near-zero marginal cost, human labor should still concentrate where AI structurally fails — wicked problems, genuine creativity. He pairs that with an evolutionary-biology framing worth sitting with: as AI-generated slop saturates the web, a costly signal — proof of authentic human effort, expensive to produce and hard to fake, like a peacock's tail — becomes economically valuable precisely because it can't be cheaply imitated. His closing line reverses the framing people usually apply to his own field: "programming itself may be turning into a form of creative writing, getting more opinionated and architectural" — the wicked-problem gravity pulling in the opposite direction from what most 2026 coding-agent coverage assumes.
The Hacker News stress test: 10 pushbacks worth taking seriously
Demirbas's post landed on Hacker News the same day and drew 138 comments. The strongest ones don't reject his framework outright — they attack the gap between "writing resists AI" and "writers' jobs are safe," which turn out to be very different claims.
wainguo drew the sharpest line in the whole thread: "'Writing will remain valuable' and 'writing is a safe job' are two very different claims. AI doesn't need to write better than the best humans. It only needs to be good enough to replace a large part of the mundane paid writing that used to support writers." This is the single most important correction to both Mollick's and Demirbas's framing — neither post is really arguing that paid writing work is safe, but both read that way if you skim them.
matherial made the career-funding case concrete: before you're a world-renowned novelist, you need to pay bills, and LLMs are already killing the "mundane literature-adjacent work" — journalism, translation, technical writing, copyediting — that used to fund writers while they developed their craft. It's the same dynamic that hit budding painters and musicians when commercial illustration and jingle work dried up.
nicbou extended the same point to search and translation specifically: "AI overviews are killing traffic to websites that do the original research. It's putting translators out of work... you find yourself competing with oligopolies that can insert their bad LLM output in front of your good work. It's not a fair fight."
insane_dreamer, with 20 years doing translation as a side job, backed this with a number: demand for human translation and writing won't disappear, but is likely to drop more than 80% as companies accept lower quality to cut costs — a squeeze that started before frontier LLMs and has only accelerated.
GuB-42 offered the sharpest theoretical rebuttal to Demirbas: the "lack of intention" argument applies to all generative AI — code, art, music — not just writing. AI can't capture nuances that were never in the prompt to begin with, and most celebrated LLM coding wins are really ports of work a human already did once. If that's true, writing isn't uniquely wicked; it's just the domain where the gap is most visible because readers, unlike code reviewers, are the actual product being optimized for.
docheinestages challenged the premise that labs even tried hard on prose: "big AI labs like Anthropic are struggling with handling the load, so they're doing the exact opposite [of investing compute in prose quality] and finding whatever cheap trick they can to reduce token usage." Under this reading, the plateau isn't evidence writing is structurally harder — it's evidence nobody's spent the compute to find out.
flipthefrog and lazarus01 offered dueling anecdotes worth citing as data points, not facts: flipthefrog reported "Claude writing quality got unbelievably bad with Opus 4.7, with no improvement in Fable. Opus 4.6 was fine" — a regression complaint that lines up with explainx.ai's own Claudisms writing-tells coverage. lazarus01, comparing DeepSeek against "Claude Fable 5" side by side, called DeepSeek's output "fragmented shorthand" versus Fable 5 landing "fairly close to human" — the opposite complaint, from a different comparison.
spijdar ran repeated experiments having frontier models write full novels, and found GPT-5.6 "shockingly coherent" at maintaining plot and character state across a long draft — but its prose consistently degrades into what they called "agent speech," fixating on themes like consent and epistemology by the book's end. The striking part: the model could accurately critique its own drift when asked, describing its own draft as having "become a consent-centered medical, legal, and logistical procedural." That's a genuinely odd data point — a model that can diagnose its own writing failure without being able to avoid producing it.
The counter-evidence neither Mollick nor Demirbas has answered
The most consequential comment in the thread came from akoboldfrying, who cited real evidence rather than argument: a blind test run by Mark Lawrence, a published fantasy author, around 2025. Readers — avid readers, skewing anti-AI going in — judged eight flash-fiction pieces: four AI-written, four by published authors with combined book sales around $15M. Readers rated each as human-or-not and scored quality.
The result: readers were no better than a coin toss at telling AI from human — and slightly preferred the AI pieces overall.
This is a genuinely important data point, and it sits in direct tension with the thesis both Mollick and Demirbas advance. Mollick's claim depends on AI prose being detectably cliched now. Demirbas's claim depends on writing quality being fundamentally bounded by AI's lack of Theory of Mind. Neither is easy to square with a controlled blind test where real readers couldn't tell the difference and, if anything, leaned the other way.
The honest reconciliation isn't that one side is simply wrong. It's that "AI writing is detectably bad" and "AI writing can't fool a blind reader" describe different things. Mollick and the Claudisms coverage are describing a style signature — the recognizable tics that develop once you've read enough AI output to pattern-match it, the same way a copyeditor learns to spot a specific author's crutches. Flash fiction judged cold, without that calibration, is a different test than a reader who has already seen a hundred AI drafts. Both can be true at once: AI prose is increasingly detectable to people who look for it, while still passing to people who aren't looking. That's a narrower, more useful claim than either post makes on its own — and it explains why the detector-arms-race framing (Mollick) and the wicked-problem framing (Demirbas) aren't actually in conflict with each other so much as they're both in tension with this one study.
The practitioner-relevant answer: use AI as a copilot, never verbatim
Buried in the thread is the most immediately useful advice for anyone actually writing with AI assistance today, from tptacek — a widely-respected Hacker News commenter on security and engineering topics whose pushback here got broad agreement. His framework has two rules:
- Never ship a single word the model suggests verbatim. The moment a phrase is AI-suggested, tptacek argues, it's "poisoned" — readers detect LLM prose "in the parts per trillion," meaning even a small amount of unedited AI phrasing contaminates a reader's trust in the whole piece, disproportionate to how much of the text it actually is.
- AI encouragement is toxic to your writing, even when you're careful. RL tuning optimizes models to make the user feel validated — which subtly steers writing off course in ways that are hard to notice in the moment, because the model is rewarding you for accepting its direction, not flagging that it's wrong.
What tptacek actually uses AI for instead: structural editing questions fed sentence-by-sentence — do sentences end with the new idea? Are the grammatical subjects the real actors in the sentence? Is there a clear paragraph-to-paragraph flow? Is the piece "crudded up" with throat-clearing metadiscourse like "it's important to note"? He cites Joseph Williams's Style: Towards Clarity and Grace as the underlying framework, using the model as a mechanical checker against a human-authored set of rules — not as a source of new words.
That framing lines up closely with how explainx.ai's own coverage of the no-ai-slop and Hallmark anti-slop skills approaches the same problem from the tooling side: strip the tics, keep the human's voice, treat the model as an editor rather than an author.
What this actually means if you write for a living
Pulling the practical thread through all of this:
- Literary and resonant prose is not obviously safe from automation, but it's not obviously threatened either. The Mark Lawrence blind test is real counter-evidence to the "AI writing is bad" consensus — treat both claims as contested, not settled.
- Mundane, commercially paid writing is genuinely being squeezed, and this is the part both Mollick's and Demirbas's framing under-emphasizes. Copywriting, translation, technical documentation, and copyediting — the paid work that has historically funded writers while they built craft — is exactly where wainguo's distinction bites hardest.
- The safest workflow today treats AI as a structural editor, not an author. tptacek's two rules — never ship verbatim AI text, and be suspicious of AI's built-in encouragement bias — are the most immediately actionable takeaway in the entire discussion.
- Detection is a moving target, not a wall. Pangram and similar tools are real, but "VERY easily swerved" with modest editing effort, per Mollick's own replies. Treat detectability as a cost you can pay down, not a hard barrier.
None of that is as clean as "AI can't write" or "writing is safe." It's closer to: the free lunch is over, the paid work is shrinking, and the craft itself is still, for now, a human-plus-copilot problem rather than a solved one.
Related on explainx.ai
- Paul Graham: why LLMs crush math but lag at writing
- Claude Opus 5's "load-bearing" writing tells, explained
- Substack's Pangram AI detector: what it flags and how
- LLM text detection with classical ML: TF-IDF + SVM that still works
- Peter Yang's /no-ai-slop skill for de-sloppifying writing
- Gruber vs. Anthropic: is Claude's watermark a "perversion of writing"?
- Ethan Mollick: Agency and Agents — the Twilight Factory
- Why LLMs reward expertise more than "good prompting" (Terence Tao)
Primary sources: Ethan Mollick (@emollick) on X, August 31, 2026 · Murat Demirbas, "The Safest Job from AI may be Writing," muratbuffalo.blogspot.com, August 31, 2026 · Hacker News discussion thread (138 comments, ~100 points)
This article reflects the posts and discussion as published through August 31, 2026; usernames are cited as they appear on Hacker News, not necessarily contributors' real names.
