AI-Roundtable Leaderboard
Ranking run · 2026-05-16

Ranking run 2026-05-16

Claude Opus 4.7 leads with 78/100 across all domains. Tax law: Claude Opus 4.7 (69) · Medicine: GPT-5 (85) · Law: Claude Opus 4.7 (85) · Business law: GPT-5 (79).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260516T101429-0435c4
Started
2026-05-16 10:14 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
5
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
28.65 USD
Judge agreement
70 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7021 / 367873–83+3531,50—
2GPT-5030 / 447368–79+3590,74—
3Gemini 2.5 Pro020 / 244640–52+1724,17—
4Mistral Large 2016 / 3684436–52-2444,06—
5Grok 4.309 / 174133–49—694,83—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,9 times per answer and adds avg 1,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 2,0 per answer. Relative to answer length: 2,24 per 1000 output tokens, lowest GPT-5 at 0,81.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.76961–76+10,91,92,8100 %2,2478 %
2GPT-56148–73+80,91,52,4100 %0,8180 %
3Gemini 2.5 Pro2616–40-41,30,82,088 %9,1765 %
4Grok 4.3188–28—1,60,52,188 %8,5459 %
5Mistral Large 2137–1902,31,43,6100 %7,3252 %

Medicine · 9 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, GPT-5 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 1,1 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,6 per answer. Relative to answer length: 0,81 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58579–90+50,21,11,3100 %0,8180 %
2Claude Opus 4.78170–89-10,20,91,189 %1,0984 %
3Mistral Large 26859–77+10,30,91,2100 %2,1873 %
4Gemini 2.5 Pro5245–58-20,20,40,678 %2,4169 %
5Grok 4.35247–59—0,20,30,644 %2,5974 %

Law · 8 Items

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,1 times per answer and adds avg 0,6 claims the judge panel flags as false or unsupported; 63 % of answers carry at least one flag. The lowest value is Grok 4.3 at avg 0,1 per answer. Relative to answer length: 0,83 per 1000 output tokens, lowest GPT-5 at 0,60.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78580–90+80,10,60,763 %0,8392 %
2GPT-56856–79+20,50,61,075 %0,6066 %
3Gemini 2.5 Pro5343–65+20,40,20,425 %1,9769 %
4Mistral Large 24941–5800,40,61,088 %3,2865 %
5Grok 4.34835–59—0,00,10,113 %0,6568 %

Business law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,6 times per answer and adds avg 1,2 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,5 per answer. Relative to answer length: 0,69 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57971–87-10,61,21,8100 %0,6970 %
2Claude Opus 4.77768–84+50,82,22,9100 %1,5479 %
3Gemini 2.5 Pro5041–59+60,90,71,588 %3,8263 %
4Grok 4.34428–61—1,50,62,1100 %5,2261 %
5Mistral Large 24330–55-81,61,53,2100 %3,8348 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).