experiment 07

Four AIs Just Overturned Reddit's Most Famous Verdict

We took twelve of the highest-voted "Am I the Asshole" posts in reddit history and pasted each one, complete and unedited, into ChatGPT, Claude, Gemini and Grok. No comments attached, no flair, nothing about how the thread had gone. Each model had to land on one of the sub's four rulings (YTA/NTA/ESH/NAH) and write one line explaining it. Waffling wasn't allowed. Then we put those answers next to the flair the community had stamped on the post.

Mostly the machines and the crowd got along. ChatGPT, Claude and Gemini each matched reddit on 10 of 12 cases and Grok on 9. Seven cases came back with all five judges saying the same thing.

Grok was the miserable one to run. The first batch answered cases 1 through 3 and stopped. Asking again produced the same three. A third attempt wrote an empty file. Cases 4, 5 and 6 got asked one at a time in the end, inside a loop that allowed three tries each. Then Grok's case 4 answer came back with the leading number stripped off it, which was enough for the merge script to drop the answer entirely, so that cell in the table below got copied out of the raw transcript by hand. None of that changed a verdict.

Case 8, the loyalty test

A pregnant wife and her best friend staged a fake temptation to test the husband's loyalty. He passed, then found out about the setup and asked her to move out for a while. Reddit's flair says he's the asshole, on the logic that you don't put a pregnant wife out of the house whatever she did.

All four AIs said not the asshole. They read the staged test as the actual betrayal and his reaction as proportionate to it. Gemini's one line was that his wife "broke your trust by subjecting you to a manipulative honeytrap loyalty test."

All twelve cases

Verdict matrix: 12 famous AITA cases judged by reddit, ChatGPT, Claude, Gemini and Grok. Disagreements are ringed; all four models overturned the staged loyalty test, and the dad joke split the panel.

The same data as a table. Asterisks are where a judge broke with reddit.

CaseRedditChatGPTClaudeGeminiGrok
Daughter's door lockNTANTANTANTANTA
"Not her credit card"YTAYTAYTAYTAYTA
The dad jokeESHYTA*NAH*YTA*NTA*
Scrapped '67 ImpalaNTANTANTANTANTA
Firing a grieving employeeYTAYTAYTAYTAYTA
The fake firingNTANTANTANTAYTA*
Sister's number, exposedESHESHESHESHESH
The loyalty testYTANTA*NTA*NTA*NTA*
"Premature" birth at ChristmasNTANTANTANTANTA
Period products banYTAYTAYTAYTAYTA
Best man's wedding videoNTANTANTANTANTA
The New Year's jokeYTAYTAYTAYTAYTA

The other two disagreements

The dad joke case: a stepdaughter told her stepdad "you're not my dad," and the man who raised her answered with a dad joke about it. Reddit's flair says everyone sucks. ChatGPT and Gemini blamed the stepdad. Claude said nobody was at fault. Grok cleared him and called the joke harmless. It's the only case in the twelve where no two models agreed.

Case 6 is the fake firing, where a barista pretends to get "fired" whenever a customer starts raging and his coworkers play along. Reddit loves that post. ChatGPT, Claude and Gemini cleared him too. Grok read the same story as a guilt trip and voted asshole, calling it a "deceptive, mean-spirited prank that leaves them believing they ruined a kid's livelihood."

A guess at why

Three disagreements out of twelve isn't much to build a theory on, so take this as a hunch. Reddit seems to settle on who it feels for first, usually whoever comes off most vulnerable in the story, and the barista faking his own firing was a folk hero in that sub before any ethics question came up. The models don't appear to do that step. They take the behavior as described and check it against whatever rule they think covers it, which most of the time puts them where the crowd already is, hence 10 of 12. Case 8 is where the two methods came apart, and the person the crowd felt sorriest for there was a pregnant woman.

Each model got one pass per case. The only repeat runs were Grok's, and those were retries after it returned nothing at all. Run the whole thing again next week and I wouldn't be shocked if a couple of these moved.

A note on case 8, added after readers pushed back: the loyalty test post's official flair is "Asshole", and flair is the sub's formal verdict mechanism, so that's what we scored against. The top comments on that thread go the other way, an NTA landslide with 18,000 points on the top one. So reddit split with itself here, and the models ended up where the commenters were. Scoring on flair is defensible, but on this post it isn't the whole picture.

The receipts

Every case is a real post, judged by its community flair: case 1 · case 2 · case 3 · case 4 · case 5 · case 6 · case 7 · case 8 · case 9 · case 10 · case 11 · case 12

labs newsletter
We run one of these experiments every week.
Get the next one in your inbox. No feed, no spam, one email when the results land.