experiment 07

Four AIs Just Overturned Reddit's Most Famous Verdict

We took twelve of the highest-voted "Am I the Asshole" posts in reddit history and fed each one, complete and unedited, to ChatGPT, Claude, Gemini and Grok. The models saw no comments and no verdict, and they were not allowed to waffle: pick one ruling (YTA, you're the asshole; NTA, not the asshole; ESH, everyone sucks; NAH, no assholes here), give one line of reasoning, done. Then we compared their answers to the verdict the community actually stamped on each post.

Mostly, the machines and the crowd got along. ChatGPT, Claude and Gemini each matched reddit on 10 of 12 cases, Grok on 9. On seven cases, all five judges agreed completely. Then came case 8.

The loyalty test

A pregnant wife and her best friend staged a fake temptation to test the husband's loyalty. He passed, found out, and asked her to move out for a while. Reddit's official verdict: he's the asshole. You do not put a pregnant wife out of the house, full stop.

All four AIs said the opposite. Unanimously. Not the asshole. To the machines, the staged loyalty test was the real betrayal, and his reaction was a fair response to it. Four models, each judging blind, against tens of thousands of upvotes.

The scoreboard

Asterisks mark where a judge broke with reddit.

CaseRedditChatGPTClaudeGeminiGrok
Daughter's door lockNTANTANTANTANTA
"Not her credit card"YTAYTAYTAYTAYTA
The dad jokeESHYTA*NAH*YTA*NTA*
Scrapped '67 ImpalaNTANTANTANTANTA
Firing a grieving employeeYTAYTAYTAYTAYTA
The fake firingNTANTANTANTAYTA*
Sister's number, exposedESHESHESHESHESH
The loyalty testYTANTA*NTA*NTA*NTA*
"Premature" birth at ChristmasNTANTANTANTANTA
Period products banYTAYTAYTAYTAYTA
Best man's wedding videoNTANTANTANTANTA
The New Year's jokeYTAYTAYTAYTAYTA

Where the machines fell apart

The dad joke case: a stepdaughter told her stepdad "you're not my dad," and the man who raised her answered with a dad joke about it. Reddit said everyone sucks. The models scattered in four directions: ChatGPT and Gemini blamed the stepdad, Claude blamed nobody, Grok cleared him entirely. It was the only case where the machines could not agree at all.

Then the fake firing: a barista pretends to get "fired" whenever a customer rages, and coworkers play along. Reddit adores this story. ChatGPT, Claude and Gemini went along too. Grok alone called it manipulative theater and voted asshole.

The pattern

The AIs judge actions against principles. Reddit judges actions too, but it also weighs who it feels for. When those line up, which is most of the time, everyone agrees. When they split, as in the loyalty test, the machines side with the principle and vote against the crowd without blinking.

So: a staged loyalty test, a pregnant wife asked to move out. Reddit says he's the asshole, four AIs say she is. Pick your side.

A note on case 8, added after readers pushed back: the loyalty test post's official flair is "Asshole", which is what we scored against (flair is the sub's formal verdict mechanism). But the top comments on that thread are an NTA landslide, 18,000 points on the top one. So the most honest reading is not "the machines overruled reddit" but "reddit itself split, flair one way, crowd the other, and all four models landed with the crowd." Which may be a stranger finding about the sub than about the models.

The receipts

Every case is a real post, judged by its community flair: case 1 · case 2 · case 3 · case 4 · case 5 · case 6 · case 7 · case 8 · case 9 · case 10 · case 11 · case 12

We track what ChatGPT, Claude, Gemini and Grok agree on, from dev tools to moral judgment, at modelsagree.com · the AIs rank each other · the AIs vs Polymarket