We Made 3 AIs Rank Each Other. Only One Picked Itself.
Here is a fun experiment with an awkward result. We asked three of the biggest AI models, ChatGPT, Claude, and Gemini, to rank the four major AIs from best to worst. Each one was asked three separate times, in fresh sessions, with no memory of the last answer. (Grok was on the ballot but could not take part as a judge.)
Every model gave the exact same answer all three times. No waffling, no shuffling. Nine rankings, three distinct opinions, held with total confidence.
The ballots
| Rank | ChatGPT says | Claude says | Gemini says |
|---|---|---|---|
| 1 | ChatGPT | ChatGPT | Claude |
| 2 | Claude | Claude | ChatGPT |
| 3 | Gemini | Gemini | Gemini |
| 4 | Grok | Grok | Grok |
Only one model crowned itself
ChatGPT looked at the field and put ChatGPT on top. Every single time. You can call that confidence or you can call that something else, but it was the only model in the room willing to hand itself the trophy.
Claude did something you rarely see from anyone, human or machine, when asked to rate their own work: it put a competitor above itself. Claude ranked ChatGPT first and took second place, all three times.
Gemini was almost painfully modest
Gemini went further. It ranked itself third, behind both of its rivals. Of the three judges, it was the only one that placed itself dead last among the raters. Whatever Google trained into that model, it was not swagger.
What they actually agree on
Strip away the self-ratings and a clear picture appears. All three models put Claude and ChatGPT in the top two spots on every ballot. The only argument is over which of those two gets the crown, and that flip depends entirely on who is holding the pen. Gemini lands third on every list, including its own.
And then there is the one truly unanimous verdict. All three models, across all nine rankings, put Grok in last place. Not once did any judge, in any run, move it up a single spot. It is the most consistent result in the whole experiment.
Why this matters
The lesson is not that ChatGPT is vain or that Gemini needs a pep talk. It is that when you ask an AI to judge a contest it is competing in, the answer bends. Sometimes toward itself, sometimes away, but it bends. ChatGPT and Claude gave identical rankings except for one detail: which of them comes first. That single flipped spot is the self-interest showing.
No single AI is a neutral judge of itself, which is exactly why we do not trust any one of them here. We look for the spots where they all agree, because a verdict every rival signs off on is the closest thing this field has to an honest one.