AI-Roundtable Leaderboard
Ranking run · 2026-05-20

Ranking run 2026-05-20

Claude Opus 4.7 leads with 76/100 across all domains. Tax law: Claude Opus 4.7 (79) · Medicine: Claude Opus 4.7 (80) · Law: Claude Opus 4.7 (77) · Business law: GPT-5 (78).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260520T074328-4a549c
Started
2026-05-20 07:43 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
5
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.2.3
Pipeline cost
76.74 USD
Judge agreement
68 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7021 / 317673–79-3551,60—
2GPT-5030 / 477470–78-1670,70—
3Gemini 2.5 Pro022 / 265248–57-2803,44—
4Mistral Large 2012 / 684945–540494,28—
5Grok 4.3013 / 194844–51+2744,59—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 2,1 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,8 per answer. Relative to answer length: 1,92 per 1000 output tokens, lowest GPT-5 at 1,01.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77972–86+60,42,12,5100 %1,9277 %
2GPT-55440–67-61,31,22,593 %1,0165 %
3Gemini 2.5 Pro3021–40-101,40,51,887 %8,3659 %
4Grok 4.33022–37+81,50,82,393 %9,1556 %
5Mistral Large 22719–37+31,81,23,0100 %6,3255 %

Medicine · 20 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,6 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 95 % of answers carry at least one flag. The lowest value is Grok 4.3 at avg 0,4 per answer. Relative to answer length: 1,50 per 1000 output tokens, lowest GPT-5 at 0,91.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78073–87-70,60,91,595 %1,5082 %
2GPT-57969–88-30,60,71,285 %0,9181 %
3Mistral Large 26759–74+30,40,91,395 %2,5668 %
4Gemini 2.5 Pro5647–6400,20,30,550 %2,3573 %
5Grok 4.35548–61+40,30,20,470 %2,0670 %

Law · 35 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 86 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,2 per answer. Relative to answer length: 1,46 per 1000 output tokens, lowest GPT-5 at 0,60.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77770–83-60,30,91,286 %1,4683 %
2GPT-57771–82+10,20,81,083 %0,6073 %
3Gemini 2.5 Pro6256–66-10,20,10,229 %1,3172 %
4Grok 4.35045–56-40,30,20,451 %3,4168 %
5Mistral Large 24638–54-41,10,71,883 %6,1262 %

Business law · 30 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 1,0 claims the judge panel flags as false or unsupported; 93 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,1 per answer. Relative to answer length: 0,53 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57872–83+10,31,01,293 %0,5373 %
2Claude Opus 4.77267–76-10,51,72,397 %1,6066 %
3Mistral Large 25246–5900,71,52,293 %3,4056 %
4Gemini 2.5 Pro5043–56+20,60,61,183 %3,7764 %
5Grok 4.34942–56+40,70,71,477 %4,4360 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).