modelsagree.com · the lab

The Lab

We put the same question to ChatGPT, Claude, Gemini and Grok and publish what happens: moral judgment, self-awareness, sealed market predictions, and the places the models refuse to agree.

Experiments

Jul 23Four AIs Just Overturned Reddit's Most Famous Verdict12 real AITA cases, judged by every model. One verdict flipped. Jul 24Show an AI Your Fight, and You Cannot Lose ItWho judges you harder, reddit or the machines? Jul 22Four AIs vs Polymarket. Every Call Public.Sealed predictions, scored when markets resolve. Jul 22We Made 3 AIs Rank Each Other. Only One Picked Itself.Blind self-confidence, measured. Jul 21Claude Knows What It Doesn't Know. Grok Doesn't Know What It Wrote.Can a model recognize its own writing? Jul 18The Cheap Tier Disagrees With the Expensive TierSame question, different price, different answer. Jul 17The Seven Tech Employers That Fell From GraceEvery AI agrees on who fell. Jul 16I Tried to Make the Best AI Agents Go RogueSafety rails, stress-tested.

The Oracle — live predictions

Jul 24The Ceasefire Has 40 Hours Left. Four AIs Say It Holds.Unanimous YES; the live-news models are the nervous ones. Will the Fed Move? Four AIs, Full Reasoning.Unanimous HOLD. Resolves July 29. Where Does LeBron Go? Four AIs, One Answer.The models pick a jersey.
What do the AIs think of your company?Type any brand, get its AI visibility grade, A+ through F. The F cards are the fun ones.