Essays on AI, software and the shape of technical work, by Cesaire Tobias.
by Cesaire Tobias
AI seems to be held to an impossible standard. We expect every output to be flawless — citations correct, claims verified, no hallucinations — while humans make mistakes constantly without facing anything like the same scrutiny. If AI is trying to approximate humanness, isn’t fallibility part of the package?
The more I pulled at this — partly in conversation with Anthropic’s Claude — the more I think the framing was wrong. Not because the worry isn’t real, but because the premise doesn’t survive contact with how we actually judge each other. Two ways it falls apart, and both are useful to think through.
My argument went like this: AI is trying to approximate humanness, humans make mistakes, so why be surprised when AI does too? Makes sense — but I got the mechanism wrong.
AI doesn’t make mistakes because it inherited human fallibility. It makes mistakes because of how it works — statistical pattern completion over training data, with no grounded model of truth and no internal verification step. AI’s imperfections are a structural artifact of their architecture. Nobody designing these systems set out to bake in our flaws as a charming touch of authenticity.
What’s borrowed from humans is the interface, not the cognition. The natural language, conversational tone and fluency are part of the human-likeness that is the surface AI produces. Perceived humaness has nothing to do with the engine driving the errors underneath.
So “AI is fallible because it’s trying to be human” is doing some sleight of hand. The fallibility and the humanness come from completely different places. The fluent prose is a deliberate design choice. The wrongness is a side effect of how the prose gets generated. Pretending they’re the same thing — that the errors are part of the human-likeness package — gives AI’s mistakes a kind of charm they haven’t earned.
The other half of my original argument was that humans don’t get held to the same standard for getting things wrong. I now think that’s not true either — at least not where it counts.
When AI hallucinates citations, invents sources, or makes things up with confidence, people object. And they should. But this isn’t a higher standard than the one we apply to humans. A journalist who invents quotes or an academic who fabricates data gets drummed out of the field. A consultant who cites studies that don’t exist becomes unhireable. Lying is treated as lying, regardless of the source — and whether the lie is delivered confidently or hedged with qualifiers doesn’t really change the verdict.
So on fabrication, AI is held to the same standard as humans, not a higher one. We don’t tolerate fabrication from each other, and we shouldn’t tolerate it from a tool that produces text at scale either.
The “humans make mistakes too” defense really only works for honest errors — typos, miscalculations, things you misremembered, points you got slightly wrong on the way to something mostly right. Those, both humans and AI get a degree of latitude on. But the loud version of the AI-criticism conversation isn’t about typos. It’s about confidently invented information presented as fact. And that’s been a fireable offense for humans for a long time.
What’s changed isn’t the standard — it’s the volume. LLMs make publishable-looking content cheap to produce, so much more of it gets shipped, and a lot of it without verification.
Some of that comes from a category mistake. We’ve inherited a prior — machines are precise — that we then misapply. Type 1200 x 32 into a calculator and the answer is exact every time. Ask a person the same and you’d expect some hesitation, an approximate guess, possibly something to double-check. AI output looks as confident as the calculator’s, but it isn’t grounded the same way. So people who would normally verify a claim before making it skip the verification when the source is AI. They aren’t being fraudulent. They’re trusting an output that looks computed when it’s actually generated.
The original premise doesn’t really hold. AI’s mistakes aren’t borrowed humanness, and on fabrication the standard isn’t higher than what we apply to ourselves. So when AI gets judged harshly for getting things wrong, that’s mostly fair — and it’s mostly symmetric.
Which leaves the question of where the asymmetry actually is. Because there is one. It just isn’t where I thought it was. The “this reads as AI” reflex, the pattern-matching against em-dashes and cadence and structure, the texture of the writing rather than its substance — that’s where I think the gap lives. That’s the argument of the companion piece, The Average Human Problem: Why AI “Sounds Like AI”.
Both threads end up pointing at the same underlying concern: whether the author engaged with what they shipped. A fabricated citation is a writer who didn’t verify. A piece flagged as “obviously AI” is often a reader detecting absence of thought. AI didn’t lower the bar — it just added a way to fail it.
The version of this I wanted to write was the one where AI is unfairly held to a higher bar. The version I came out with is more deflating but probably more useful: the bar isn’t really about AI at all. It’s about whether someone took responsibility for what came out. The interesting question isn’t why we judge AI harshly. It’s why we sometimes pretend we judge each other any more gently.
May 4, 2026
tags: ai