Which AI should you actually use?
With Whitney Lee and Michael Forrester
Which AI should you actually use? Answered in three solo segments with every source funding-flagged.
In this episode · 3 segments
ChatGPT, Gemini, Claude, Grok: is there a fair way to compare them?
Evidence · STRONG. Leaderboard Illusion (2025) [arXiv][Mixed: Cohere-affiliated]: private-variant gaming across 2M LMArena battles.
Read the transcript
Is there a fair way to compare ChatGPT, Gemini, Claude, and Grok? This episode asks one question: which AI should you actually use? And this first segment asks whether the rankings everyone quotes can answer it.
The most famous ranking is LMArena, formerly Chatbot Arena. The idea is lovely. Two anonymous models answer your prompt side by side, you vote for the better answer, and the votes get crunched into an Elo rating, the same kind of rating chess players carry. Crowds of real people voting blind. It feels like the fairest test we have.
In April 2025, researchers from Cohere Labs, Princeton, Stanford, and MIT released a paper called "The Leaderboard Illusion," since peer reviewed at NeurIPS 2025. Funding flag, out loud: Cohere builds and sells models too, so this critique comes from a company with its own horse in the race. Keep that in your pocket while I walk you through what they documented across roughly two million arena battles.
First finding: undisclosed private testing. Before Meta released Llama 4, it quietly tested 27 different variants on the arena. Test a pile of variants, keep the score you like, and the public sees one polished number. Picture a drug company running 27 trials and publishing the winner. The study also documented selective retraction of scores.
Second finding: the prompts themselves. Up to 26.5 percent of arena prompts were duplicates or near duplicates. A model tuned to ace the questions people ask over and over climbs the board without getting better at anything new.
Third finding, the one that stings. The team fine-tuned a model on arena-style data and its win rate on ArenaHard jumped 112 percent. Its score on MMLU, a standard knowledge benchmark, did not rise at all. It slipped slightly. You can climb the most famous leaderboard in AI while becoming no more capable.
Fairness requires the other side. LMArena published a rebuttal disputing the framing, and per our evidence standard I will tell you to read both. Also on the record: LMArena is now VC-backed, per our evidence standard review as of July 2026. The referee has investors.
The rating math is shaky on its own terms. One 2025 arXiv preprint, not yet peer reviewed, showed that dropping just a handful of preference votes can change which model sits at number one. A second 2025 preprint showed that a few hundred deliberately rigged votes, out of roughly 1.7 million, can move a target model's rank on purpose. When the top spots are that close together, the ordering is noise.
And suppose every vote were clean. What does an arena vote measure? It measures which answer a stranger preferred in the moment: style, confidence, formatting. Our evidence standard files LMArena under human preference for exactly that reason. Preference is worth knowing. Capability is a different question, and a vote does not settle it.
So, is there a fair way to compare these four chatbots by leaderboard? Here is the grade. The evidence is strong that the most popular comparison surface can be gamed and has been gamed: a peer-reviewed forensic study, two independent preprints on rating fragility, and a documented case of climbing the board with zero capability gain.
What should you do? When a headline says one chatbot beat another by a couple of Elo points, treat that gap as marketing. And before you trust any leaderboard score, ask how many private variants ran before that number went public.
Sources
- The Leaderboard Illusion. Singh, Nan, Wang, et al. (Cohere Labs, Princeton, Stanford, MIT). NeurIPS 2025, peer reviewed; preprint April 2025. [Mixed: Cohere-affiliated]. https://arxiv.org/abs/2504.20879
- Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings. arXiv preprint, 2025. [No funding flag recorded in research notes]. https://arxiv.org/pdf/2508.11847
- Improving Your Model Ranking on Chatbot Arena by Vote Rigging. arXiv:2501.17858, 2025. [No funding flag recorded in research notes]. https://arxiv.org/html/2501.17858
- The AI Inevitable Evidence Standard v1.1 (July 2026): LMArena evaluator entry (human-preference scope; VC-backed caveat). Internal: docs/research/evidence-standard.md
What does the science say about which AI is best?
Evidence · STRONG. NeurIPS 2025 construct-validity review [PR][Independent]: benchmark scores do not license general-capability claims; the question is malformed.
Read the transcript
Whitney just showed you the arena problem. Here's what the science says about the deeper issue. The question sounds simple: which AI is best? Line up ChatGPT, Gemini, Claude, and Grok, test them, crown a winner. I want to show you why the measurement science says that question, asked that way, has no answer.
Start with the biggest study we have. In November 2025, a paper in the NeurIPS Datasets and Benchmarks track, more than 40 authors working with 29 expert reviewers, examined 445 large language model benchmarks from the top NLP and machine learning venues. Funding flag: independent, and peer reviewed. Their finding: pervasive construct-validity problems. Construct validity is the measurement question: does this test measure the thing it claims to measure? Across those 445 benchmarks, the answer was, frequently, no. The paper closes with eight corrective recommendations. That is a structural indictment of how the field measures itself.
This finding has roots. Back in 2021, Raji and colleagues called the framing of benchmarks as tests of general capability "dangerous and deceptive," and Bowman and Dahl laid out validity criteria that most NLP benchmarks fail. A February 2025 interdisciplinary review on arXiv traces today's problems straight back to those two papers. The field has known for five years. The 2025 audit confirmed it at scale.
Now watch it happen to one concrete benchmark. SWE-bench is the flagship test for AI coding: can the model's patch resolve a real software issue. Wang, Pradel, and Liu, peer reviewed at ICSE 2026, with no vendor flag recorded in our research files, found that limited test coverage lets about 7.8 percent of merely plausible patches count as correct even though they fail the full developer test suite. That inflates reported resolution rates by about 6.4 points. A second study, UTBoost, an arXiv preprint from June 2025, strengthened the test suites and watched the verdicts move: 24.4 percent of SWE-bench Verified entries were affected, 40.9 percent on SWE-bench Lite, and the leaderboards reshuffled with 18 ranking changes on one board and 11 on the other. Better tests produced different winners on the same benchmark.
Two more problems stack on top. Saturation first: MMLU and MMLU-Pro scores for frontier models sit above roughly 88 percent, per a 2026 industry report from Kili Technology. Flag that source out loud: an industry report, directional only. When every frontier model scores 88-plus, the gaps between them are statistically uninformative. Then the harness: a 2026 arXiv study measured that the software wrapper around a model, the scaffolding and prompts and tools, produced 7.8 times more variance than the choice of model itself, and per our evidence standard the same weights swing 10 to 35 points on agentic benchmarks by scaffold alone. A benchmark score is a property of the model plus its scaffold plus the harness plus the test subset. The score cannot crown "the model," because the model alone never took the test.
Here is what benchmarks are still good for, and this part matters. They separate weak models from strong ones reliably. A model that fails everything is genuinely weaker than one that passes most things. The filter works. What the science will not support is using those same scores to rank the top handful of frontier models against each other, where the gaps are small, gameable, and sitting inside measurement noise. Our research brief puts a working number on it: treat gaps under five points at the frontier as noise.
The verdict, as a grade: the evidence here is strong, and what it establishes is that "which AI is best" is a malformed question. Best at what, measured how, inside which harness? Only task-specific evaluation survives the construct-validity literature.
So use benchmarks the way the evidence licenses. Let them filter the field down to a shortlist of two or three credible models. Then settle the question on your own tasks, with your own success criteria. The next segment shows you how to run that test in an afternoon.
Sources
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks. Bean, Kearns, et al. (40+ authors, 29 reviewers). NeurIPS 2025 Datasets and Benchmarks Track. [PR][Independent]. https://arxiv.org/abs/2511.04703
- Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation (citing Raji et al. 2021 and Bowman & Dahl 2021). arXiv, Feb 2025. https://arxiv.org/html/2502.06559v1
- Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study. Wang, Pradel, Liu. arXiv:2503.15223; ICSE 2026, peer reviewed. [No vendor flag recorded in research notes]. https://arxiv.org/abs/2503.15223
- UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench. Yu, Zhu, He, Kang. arXiv:2506.09289, June 2025. [No vendor flag recorded in research notes]. https://arxiv.org/abs/2506.09289
- AI Benchmarks 2026: Top Evaluations and Their Limits. Kili Technology, 2026. [Industry report, directional only]. https://kili-technology.com/blog/ai-benchmarks-guide-the-top-evaluations-in-2026-and-why-theyre-not-enough
- Harness-variance study (2026, arXiv): harness variance 7.8x model variance; scaffold swings of 10 to 35 points. As recorded in the show question pool and Evidence Standard v1.1 (three-pass verified, July 2026). Internal: docs/research/question-pool.md, docs/research/evidence-standard.md
- Do Benchmark Scores Actually Predict Which Model Wins in Real Work? (sub-5-point noise rule; coarse-filter guidance). Internal brief: docs/research/model-choice-benchmarks-vs-execution.md
The one test that matters: your own head-to-head in an afternoon.
Evidence · STRONG methodology model: METR developer RCT (2025) [Independent]: 19% slower while feeling 20% faster.
Read the transcript
So which chatbot should you actually use? The answer comes from a test you run yourself, on your own work, in one afternoon. Michael just showed you why benchmarks can only cut the field to a shortlist; this segment settles the shortlist.
Start with the obvious objection. Why run a formal test at all? Why not just use both models for a week and trust your gut? Because your gut has been measured under controlled conditions, and it failed.
In 2025, METR, an independent research group, published a randomized controlled trial of experienced software developers doing real work with and without AI assistance. Funding flag, out loud: our research files mark this study independent, and they also note METR has had Anthropic funding ties, with transparent methods, so keep that in your pocket. The result: developers working with AI took 19 percent longer to finish than developers working without it. Then the researchers asked those same developers how it went. They believed the AI had sped them up by about 20 percent. Wrong direction, wrong size. Experienced professionals, inside a controlled trial, could not feel a 19 percent slowdown happening to them in real time.
That is the whole case for measurement. If seasoned developers cannot feel a slowdown that large, you cannot feel the difference between two chatbots either. Feelings do not survive contact with a stopwatch. So bring a stopwatch.
Here is the recipe, five steps.
Step one, pick your tasks. Go back through your last week of work and pull five to ten real tasks you did: the email you drafted, the code you debugged, the document you summarized, the analysis you built. Real tasks only, from your own week, because the whole point is to test the work you will hire the model to do. Our research brief sizes a full private eval at 20 to 100 tasks; five to ten is the afternoon version, and it already tells you more about your decision than any public leaderboard.
Step two, define done before you start. For each task, write the pass criteria before either model sees the prompt. Does the code run and pass your tests? Does the summary contain the numbers you needed? Did it catch the bug you know is in there? Verifiable checks beat preference every time; that comes straight from the brief. If you decide what counts as good after reading the answers, you are voting on style, and you already know where voting on style leads.
Step three, run the head-to-head. Same task, same prompt, both contenders. And time every task, start to finish, including your cleanup of the model's output. The METR result is the reason the timer is mandatory. Time-to-done is exactly the number your feelings will fake.
Step four, score blind where possible. Put the outputs side by side with the model names stripped, then grade against your written criteria. Writing the criteria in advance blinded the judging; stripping the names finishes the job.
Step five, count. Wins, losses, and total time per model. If one model wins clearly, you have your answer, backed by the only evidence that applies to you. If the result is close, the honest conclusion is that these two are interchangeable for your work, so pick on price.
Two upgrades from the playbook. If you plan to trust a task unattended, run that task several times, because a model that succeeds once has told you nothing about succeeding reliably; measure at the reliability you need. And when either model ships an update, re-run your set, because rankings churn.
Now the episode verdict, as a grade. The evidence is strong, across all three segments, that a fair universal comparison of these chatbots does not exist. The most famous leaderboard has been gamed. The benchmark literature says "which AI is best" is a malformed question, because no public score can certify general capability. Construct validity is whether a test measures what it claims to measure, and exactly one benchmark has construct validity for your decision: the one built from your own tasks, graded by your own definition of done.
So here is the close for the whole episode. Use the public numbers once, as a coarse filter: treat gaps under five points between frontier models as noise, and cut the field to two contenders. Then spend the afternoon. Five to ten real tasks, pass criteria written first, a timer running, names stripped before you grade. When the models update, run it again.
Sources
- METR developer RCT: experienced developers 19% slower with AI assistance while estimating themselves about 20% faster. arXiv, 2025. [Independent; METR's Anthropic funding ties noted in the internal brief]. As recorded in docs/research/question-pool.md (cluster T3).
- Do Benchmark Scores Actually Predict Which Model Wins in Real Work? (private-eval playbook: 20 to 100 task set; verifiable checks over preference; sub-5-point noise rule; reliability-bar and re-test guidance; METR ties caveat). Internal brief: docs/research/model-choice-benchmarks-vs-execution.md
- The Question Pool, cluster T3 evidence states (verified 2026-07-04). Internal: docs/research/question-pool.md
Topics