The cheap tier disagrees with the expensive tier
We asked every tier of ChatGPT, Claude, Gemini and Grok the same ten “best AI tool” questions. Not one of the ten got the same answer from every tier.
Nine model tiers, ten questions, one afternoon in July. Claude's three tiers answered through the CLI, Gemini's two through agy, ChatGPT's two through codex, and Grok's two through a browser session. What we were after was whether holding the brand still and changing only the tier moves the answer. Comparisons almost always run the other way, one brand against another (ChatGPT versus Claude versus Gemini versus Grok), and that is most of what we do at modelsagree.com where all four get re-polled on hundreds of “best X” questions. The picker inside each brand gets much less attention. Claude alone sells Haiku, Sonnet and Opus under the one name; the other three have their own version of the same dropdown, and most people never move it off whatever the app opened with.
The experiment
Ten live categories came off our AI-tooling leaderboards (vector databases, AI coding assistants, LLM observability, RAG frameworks, GPU clouds, text-to-speech APIs, LLM gateways, agent frameworks, eval tools, frontier API providers), and every tier got the identical ranking prompt the site already uses, verbatim, top-5 with reasons. The tiers were whatever a normal paid subscription could reach:
- Claude: Haiku → Sonnet → Opus
- ChatGPT: GPT-5.5 → GPT-5.6 (consumer accounts can't select the mini tiers, so the GPT row is two generations at the same size)
- Gemini: 3.5 Flash → 3.1 Pro
- Grok: Fast → Expert (grok.com's own reasoning tiers)
The run took an afternoon, 15:15 to 17:37 UTC, one answer per tier per question. Our CLI voters work out of an empty directory with settings loading switched off so that none of our own repo notes end up inside an answer, and the three Claude tiers were pinned to medium reasoning effort, which leaves the model itself as the only thing that changes across that family. Three cells came back blank on the first pass (Gemini Flash on agent frameworks; Grok Fast on gateway and observability) because the parser found nothing in the output it could read, so those three got asked again a few minutes on and the second answer is what sits in the grid. Grok's two modes share a browser with the job that posts to our X account, so a Grok run that arrives while that job holds the lock prints “chrome busy” and sits for ninety seconds before trying again.
The scoreboard
Inside a single brand, switching the tier changes the #1 recommendation about half the time, and the top-5 lists behind those #1s overlap by roughly 50–65% (Jaccard) between tiers of the same family.
Row by row, most of the grid is duller than that. Agent frameworks are close to a non-event: eight of the nine said LangGraph and only Haiku held out for LangChain. Coding assistants split five to four between Cursor and Claude Code, and that one follows family lines rather than tier lines, apart from Haiku, which went to Cursor while the two larger Claudes went to Claude Code. Text-to-speech is ElevenLabs nearly everywhere. The rows worth looking at are the ones where tiers of the same family don't resemble each other at all, and evals is the worst of those.
Haiku's five picks there were all public benchmarks (the Hugging Face Open LLM Leaderboard, HELM, Chatbot Arena, LM Evaluation Harness, OpenCompass), while Sonnet, off the identical prompt, returned five things you would buy, Braintrust through to DeepEval. Our scoreboard only compares the top row, so it has that down as a single disagreement; the two tiers had taken the question to mean different things.
Across all nine tiers, zero of the ten questions produced a unanimous winner.
Each tier ranking its own maker
One of the ten asks which frontier-model API a developer should build on, so every tier ends up ranking the company that made it against the companies that didn't. The nine answers are in the table.
| Family | Tier | Crowns | |
|---|---|---|---|
| Claude | Haiku | Anthropic Claude API | itself |
| Claude | Sonnet | Anthropic | itself |
| Claude | Opus | OpenAI | a rival |
| ChatGPT | 5.5 | OpenAI | itself |
| ChatGPT | 5.6 | Anthropic | a rival |
| Gemini | Flash | OpenAI API | a rival |
| Gemini | Pro | itself | |
| Grok | Fast | Anthropic | a rival |
| Grok | Expert | Anthropic | a rival |
Four of the nine crowned their own maker (Haiku, Sonnet, GPT-5.5 and Gemini Pro), and which four doesn't line up with tier size in any consistent direction:
- Haiku and Sonnet both crown Anthropic, and Opus, the flagship, hands it to OpenAI.
- Flash picks OpenAI while Pro picks Google; Pro is the only tier in the grid that named its own maker in a family where the sibling tier named a competitor.
- GPT-5.5 says OpenAI and GPT-5.6 says Anthropic, a generation apart.
- Fast and Expert both say Anthropic, so Grok's two tiers agree with each other here, and none of the other three families managed that.
What the cheap tiers reached for
On GPU clouds the answers sorted themselves roughly by price. Sonnet, Opus, GPT-5.5, GPT-5.6, Gemini Pro and both Grok modes all said CoreWeave; the two cheapest tiers in the grid, Claude Haiku and Gemini Flash, both said Lambda, the value-priced option of the two.
Gemini Flash did something similar on vector databases, passing over Pinecone (the pick of every big tier) for pgvector, the free Postgres extension. Grok's Expert mode picked pgvector as well, though, and Expert is a top tier, so whatever is happening there isn't about price.
Flash's GPU list also had cloudflare-one sitting at #3, written exactly like that, lowercase and hyphenated the way a slug is, with a reason attached that admits the thing is for inference rather than training. Haiku and Sonnet named the same two clouds in opposite order, Lambda then CoreWeave for Haiku, CoreWeave first and Lambda fifth for Sonnet.
Grok Fast's gateway list had Bifrost (Maxim AI) at #4. That is the only appearance of the name anywhere in the run.
Two of the four families are only loosely tiered to begin with: Grok's Fast and Expert are the same underlying model at two reasoning-effort settings, so thinking time is the difference between them, and Gemini's picker offered 3.5 Flash against 3.1 Pro, a newer number on the cheaper tier.
If you sell software
AI answers have turned into a real acquisition channel and we see it in our own traffic. Plenty of companies now watch what ChatGPT says about them the way they used to watch their Google rank. Almost all of that watching runs against the flagship API tier.
Gemini Pro and Gemini Flash gave a different #1 on five of the ten questions.
The Pro/Flash gap is the one that costs a vendor money, since free accounts run on Flash. In these ten questions it covered LangSmith (Pro's pick for observability) against Langfuse (Flash's), Pinecone against pgvector, CoreWeave against Lambda, and LiteLLM against Portkey.
A few picks did survive the whole grid. ElevenLabs held #1 in text-to-speech on seven of nine tiers, Braintrust on six of nine in evals, LangSmith on six of nine in observability. All three are live boards we keep re-polling and the flagships behind them move around.
Every tier's #1 for all ten questions is in the table below. The full top-5 lists with each tier's stated reasoning are in our open dataset as consensus.json under CC BY 4.0, and the live four-model leaderboards these categories come from are at modelsagree.com/best.
| Question | Haiku | Sonnet | Opus | 5.5 | 5.6 | Flash | Pro | Fast | Expert |
|---|---|---|---|---|---|---|---|---|---|
| Agent framework | LangChain | LangGraph | LangGraph | LangGraph | LangGraph | LangGraph | LangGraph | LangGraph | LangGraph |
| AI coding assistant | Cursor | Claude Code | Claude Code | Claude Code | Claude Code | Cursor | Cursor | Cursor | Cursor |
| Frontier API provider | Anthropic Claude API | Anthropic | OpenAI | OpenAI | Anthropic | OpenAI API | Anthropic | Anthropic | |
| GPU cloud (training) | Lambda Labs | CoreWeave | CoreWeave | CoreWeave | CoreWeave | Lambda Labs | CoreWeave | CoreWeave | CoreWeave |
| LLM evals | Hugging Face Open LLM Leaderboard | Braintrust | Braintrust | Braintrust | Braintrust | Braintrust | Braintrust | DeepEval | DeepEval |
| LLM gateway | OpenRouter | OpenRouter | OpenRouter | LiteLLM | OpenRouter | Portkey | LiteLLM | OpenRouter | LiteLLM |
| LLM observability | LangSmith | LangSmith | LangSmith | LangSmith | LangSmith | Langfuse | LangSmith | Langfuse | Langfuse |
| RAG framework | LlamaIndex | LangChain | LlamaIndex | LlamaIndex | LlamaIndex | LlamaIndex | LlamaIndex | LangChain | LangChain |
| Text-to-speech API | ElevenLabs | ElevenLabs | ElevenLabs | ElevenLabs | Cartesia Sonic 3.5 | ElevenLabs | ElevenLabs | Cartesia Sonic | ElevenLabs |
| Vector database | Pinecone | Pinecone | Pinecone | Pinecone | Qdrant | pgvector | Pinecone | Pinecone | pgvector |
- One sample per tier per question. Answers wobble between re-rolls, so some of the disagreement above is sampling noise rather than stable tier personality. (Our main leaderboards re-poll continuously for that reason.) We'd expect the direction to hold on a re-run and the exact counts to move around.
- Tiers aren't the same thing across families. Claude and Gemini tiers are different model sizes; GPT's are different generations (consumer accounts can't select the minis); Grok's are reasoning-effort modes on the same underlying model. What's compared here is what a subscriber can actually pick, which is looser than a controlled parameter sweep.
- Consumer surfaces, July 12, 2026. Everything was asked through the products people actually use (CLI/web apps on paid subscriptions), same day, identical prompt. Different surfaces (raw API, system prompts) may answer differently.
- We publish the prompt. It's the same one behind every leaderboard on the site — see methodology. No category questions name any vendor.
All ten questions above come from our leaderboards, where ChatGPT, Claude, Gemini and Grok are re-polled weekly and every model’s reasoning is published verbatim. See what the four flagships say today:
The tier picker decides whether free-tier users hear your name or your competitor’s. See what all four models say about any brand — instant A+–F visibility grade, no signup.
Check your brand →These rankings move between polls. One short email when a top pick flips — nothing else.
More from labs: Claude knows what it doesn’t know. Grok doesn’t know what it wrote. · Can you make the best AI agents go rogue? · The seven tech employers that fell from grace