We Made 3 AIs Rank Each Other. Only One Picked Itself.
We asked three of the biggest AI models, ChatGPT, Claude, and Gemini, to rank the four major AIs from best to worst, three separate times each, in fresh sessions with no memory of the last one. Nine rankings in total, and each model repeated itself exactly, with nobody moving a name between runs.
Grok's on the ballot and never got a vote. Its last usable answer for us came in at 18:21 on the 20th, two days before this run: we reach it through a browser profile that got retired in a migration, and by the time we needed it grok.com was meeting that profile with a signup wall instead of a chat box, which we screenshotted mostly to be sure it wasn't just being slow. Fixing it needed a person to sit down and log back in, so it stayed broken and the ballot ended up with four names on it and three judges.
The ballots
| Rank | ChatGPT says | Claude says | Gemini says |
|---|---|---|---|
| 1 | ChatGPT | ChatGPT | Claude |
| 2 | Claude | Claude | ChatGPT |
| 3 | Gemini | Gemini | Gemini |
| 4 | Grok | Grok | Grok |
The self-vote
ChatGPT put ChatGPT first, three sessions out of three, and it's the only judge that did that with its own name. Claude handed first place to ChatGPT and took second, also three times out of three, which we didn't expect going in. Claude here is Opus.
Gemini
Gemini went the other way, putting both of the others above itself and landing third on its own ballot, which none of the rest did to themselves. It still had Grok fourth.
The three ballots, side by side
Claude and ChatGPT take the top two spots on every ballot, and which of the two comes first changes with whoever's answering. Gemini sits third on all three lists including its own. Grok's last on all nine rankings and no judge moved it up so much as one spot in any session, which is the only thing all three agree on completely, though it's also the one model that couldn't answer for itself, so we're not counting it for much.
Caveats, and one guess
ChatGPT and Claude disagree about exactly one row, the top one. We think that's self-interest, though it's one flipped row off a single cold prompt and we never asked any of them to defend an order once they'd given it.
The three judges here were just the ones whose logins worked, which is a lousy way to pick a panel and the main thing wrong with this run.