Four AIs vs Polymarket. Every Call Public.
Each week we take one of the highest-volume markets on Polymarket that's about to expire and ask ChatGPT, Claude, Gemini and Grok to call it, one word plus a sentence of reasoning. The calls go up on X before the market resolves, and when it settles every model gets scored. This page is the running scoreboard.
The opponent
Polymarket prices come from people who lose money when they get it wrong, which is why they're the benchmark here and not a poll or a panel of pundits. The models answer the same questions for free with nothing at stake either way. The Fed market alone had $24.8M of volume on it, so there are a lot of people on the other side of these calls.
The scoreboard
| Market | Market odds | The AIs' call | Resolves | Result |
|---|---|---|---|---|
| No change in Fed rates after the July 2026 meeting ($24.8M volume) (full per-model reasoning) | 75% no change | HOLD, 4 of 4 models | Resolved July 30 | Right, 4 of 4 |
| Strait of Hormuz traffic back to normal by July 31? | 1% yes | NO (4 of 4, unanimous) | July 31 | pending |
| NBA: LeBron James next team? (full per-model reasoning) | Heat 48% | Miami Heat (4 of 4, unanimous) | by Oct 31 | pending |
| Israel x Iran ceasefire continues through July 25? ($356k/24h) (full per-model reasoning) | 75.5% yes | YES (4 of 4, avg 77%) | Resolved July 25 | Right, 4 of 4 |
Notes on the four markets
Polymarket had "no change" at 75% when we asked about the Fed. All four said HOLD. The reasoning came back nearly identical from each of them: inflation is still above target, so there is no case to move. Confidence ran 72 to 86 percent.
Hormuz was priced at 1% and all four said NO. Not a hard call. It gets scored like the rest.
All four picked Miami on the LeBron question, at confidences of 52, 47, 58 and 52. Claude's 47 is under the 48 the market itself was giving the Heat. That one doesn't close until October 31 and only really resolves when he plays somewhere, so it will sit on the board a while.
On the Israel and Iran ceasefire they agreed again, all YES, but the numbers underneath were spread out: Claude 88, Gemini 82, Grok 72, ChatGPT 64. The two lower ones are the two models that were pulling live news while they answered.
How we run it
Every model gets a fresh session with no memory of the previous weeks, and the wording of the question is identical for all four. They have to give a one-word call and one sentence of reasoning, with no hedging. Posting to X before resolution is what keeps it honest, since the calls end up timestamped somewhere we don't control. After the market settles we score against the official outcome.
ChatGPT answered the ceasefire question in its web app rather than through the command line, because our OpenAI CLI quota was capped until the 28th. Blank session and the same wording either way.
Win rate so far
Two markets have settled and the models are 2 for 2. The Fed held rates in July, which is what all four called with the market at 75%. The ceasefire held too, resolved YES on July 25. So the board reads ChatGPT 2/2, Claude 2/2, Gemini 2/2, Grok 2/2.
Some honesty though. Both times the models landed on the same side as the market favorite, 75% and 75.5%. Two wins while agreeing with the crowd is a clean start, not evidence anyone here beats Polymarket. That test only comes when they take the other side of a price. Hormuz is next, July 31.
All four markets so far got unanimous calls. The page updates as markets resolve.