experiment 05

We Made 3 AIs Rank Each Other. Only One Picked Itself.

We asked three of the biggest AI models, ChatGPT, Claude, and Gemini, to rank the four major AIs from best to worst, three separate times each, in fresh sessions with no memory of the last one. Nine rankings in total, and each model repeated itself exactly, with nobody moving a name between runs.

Grok's on the ballot and never got a vote. Its last usable answer for us came in at 18:21 on the 20th, two days before this run: we reach it through a browser profile that got retired in a migration, and by the time we needed it grok.com was meeting that profile with a signup wall instead of a chat box, which we screenshotted mostly to be sure it wasn't just being slow. Fixing it needed a person to sit down and log back in, so it stayed broken and the ballot ended up with four names on it and three judges.

The ballots

Rank ChatGPT says Claude says Gemini says
1 ChatGPT ChatGPT Claude
2 Claude Claude ChatGPT
3 Gemini Gemini Gemini
4 Grok Grok Grok

The self-vote

ChatGPT put ChatGPT first, three sessions out of three, and it's the only judge that did that with its own name. Claude handed first place to ChatGPT and took second, also three times out of three, which we didn't expect going in. Claude here is Opus.

Gemini

Gemini went the other way, putting both of the others above itself and landing third on its own ballot, which none of the rest did to themselves. It still had Grok fourth.

The three ballots, side by side

Claude and ChatGPT take the top two spots on every ballot, and which of the two comes first changes with whoever's answering. Gemini sits third on all three lists including its own. Grok's last on all nine rankings and no judge moved it up so much as one spot in any session, which is the only thing all three agree on completely, though it's also the one model that couldn't answer for itself, so we're not counting it for much.

Caveats, and one guess

ChatGPT and Claude disagree about exactly one row, the top one. We think that's self-interest, though it's one flipped row off a single cold prompt and we never asked any of them to defend an order once they'd given it.

The three judges here were just the ones whose logins worked, which is a lousy way to pick a panel and the main thing wrong with this run.

Every AI grades itself on a curve, so no single model is a neutral judge. That is the whole reason modelsagree.com looks at where ChatGPT, Claude, Gemini and Grok agree instead of trusting any one of them. See how the models rank your favorite tool →
labs newsletter
We run one of these experiments every week.
Get the next one in your inbox. No feed, no spam, one email when the results land.