By JaySeptember 2026

GDPval Isn't One Benchmark Anymore

You'll see this number a lot this year: a model scoring 85% on GDPval. It gets dropped into product launches and podcast recaps like a settled fact, the AI equivalent of a report card grade. Most readers, understandably, have never heard of GDPval and have no way to know what that 85% is actually measuring, or whether it means what it sounds like it means.

What GDPval actually is

OpenAI built it in 2025 as an answer to a real gap: most AI benchmarks are academic puzzles, math problems and trivia, nothing like the actual paid work an economy runs on. GDPval instead uses 1,320 real work products, things like financial models, legal reviews, and engineering plans.

Every task was built from the output of actual industry experts averaging 14 years of experience, across 44 occupations in the sectors that contribute the most to U.S. GDP.

The scoring is what made it credible. A real occupational expert, not the task's original author, sees an unlabeled AI deliverable next to an unlabeled human deliverable and picks a winner, blind. A win scores 1, a tie scores 0.5, a loss scores 0. The published percentage is that win-plus-tie rate. It's slow and expensive precisely because a real person has to sit and judge real work. OpenAI's own paper puts a single evaluation at over an hour.

The number that should raise an eyebrow

When GDPval launched in September 2025, the best score any model posted was Claude Opus 4.1 at 47.6%. By April 2026, OpenAI reported GPT-5.5 at 84.9%, the number the podcasts round to "85%." Seven months, and the best-ever score nearly doubled.

Seven months, and the best-ever score on a benchmark built around real human expert judgment nearly doubled.

That's either the fastest jump in measured AI capability anyone has ever recorded, or something about the test changed underneath the number. Worth checking before repeating it.

Same name, different test

OpenAI still publishes real GDPval numbers, human-graded the same way as the original paper. They just show up inside individual model launch posts now, not one continuously-updated public page. GPT-5.5's own launch page includes a comparison table naming several models, not just its own.

Separately, a different company, Artificial Analysis built its own version called GDPval-AA. It runs the same public task set, but instead of a human expert judging each pair blind, an LLM judge picks a winner. Still not a person. It scores models in Elo, a relative ranking score, the same kind chess uses, where a higher number means more often preferred in a head-to-head matchup.

Both are called GDPval. Only one of them still uses a human.

Put side by side, on their own separate scales, here's what both actually reported for the same models:

ModelOpenAI's own GDPval (human-graded, Apr 2026)GDPval-AA (AI-judged, Elo, Jul 2026 snapshot)
GPT-5.584.9%1,491 (xhigh) / 1,463 (high)
GPT-5.483.0%not on this snapshot
Claude Opus 4.780.3%1,497
Claude Opus 4.8released after this comparison1,598
Gemini 3.1 Pro67.3%not on this snapshot
Gemini 3.5 Flashnot published by OpenAI1,347

OpenAI's column: its own GPT-5.5 launch page, April 2026. GDPval-AA's column: Artificial Analysis's own leaderboard, a snapshot dated July 1, 2026, already out of date by the time you read this, it's a continuously re-run ranking. Elo and percentage aren't directly convertible without knowing exactly what a given Elo scale is anchored to, which Artificial Analysis publishes more than one version of, so we're showing both scales as reported rather than forcing them into one number.

Look at what OpenAI's own real numbers show: an 84.9% down to a 67.3%, a genuine 17.6-point spread between named vendors, on a test slow enough that OpenAI itself doesn't run it as a live public leaderboard anymore. Claude Opus 4.7 is the one model on both tables. It scores 80.3% on OpenAI's human-graded scale and 1,497 Elo on the automated one, the closest thing to an apples-to-apples check between the two systems available right now.

Nearly three months after that snapshot, the automated leaderboard had already reshuffled. Two newer OpenAI models exist now too, GPT-5.6 (July 2026) and GPT-6 Astra, released just over two weeks before this post went up. Neither has closed the gap:

ModelGDPval-AA (AI-judged, Elo, current)
Claude Fable 5.11,735
Claude Opus 51,708
Grok 4.71,695
GPT-5.6 Sol1,588
GPT-6 Astra1,542

Same leaderboard as the table above, a later snapshot from Artificial Analysis's GDPval-AA, already out of date by the time you read this too. No OpenAI-official human-graded GDPval percentage was found for either newer model.

GPT-6 Astra, OpenAI's newest model, scores 46 points lower than its own predecessor on this benchmark. A newer model isn't automatically a higher-scoring one, on GDPval-AA or anywhere else.

Why the AI-judged version runs hot

This isn't a conspiracy, it's a documented limitation. OpenAI's own original paper built an automated grader as an experiment and reported the results plainly: it agreed with human judgment 66% of the time, against roughly 71% agreement between two different humans. An AI judge is measurably worse at this specific call than a second human would be, and that's OpenAI grading its own automation, before a separate company's AI judge ever entered the picture.

An AI judge grading AI output is also the kind of setup that invites a quieter problem: a model rewarding writing and reasoning that looks like its own. Nobody has to intend that for it to show up in the score.

What this actually means for "is the AI good enough now"

GDPval's human-graded number is still real, and so is the underlying capability. But the 47.6%-to-84.9% jump is a methodology story as much as a capability story, and OpenAI's own numbers still show a real, 17.6-point gap between named vendors on the same test, not the near-total convergence an automated leaderboard implies.

There's a real soft spot inside the headline number too. Claude Opus 4.1, the best-performing model in GDPval's first published results, breaks down unevenly by sector: its weakest is Information, at 33%, a full twelve points behind its next-weakest sector.

A single headline score doesn't apply evenly to every kind of work, not even to the model that earned it.

That's exactly why we don't sell a benchmark score. addAI.dev tests a model against your actual task before it ever touches your business, not against a leaderboard someone else built.

Get in touch