AI-Roundtable Leaderboard
Ranking run · 2026-06-07

Ranking run 2026-06-07

Claude Opus 4.7 leads with 81/100 across all domains. Tax law: Claude Opus 4.7 (77) · Medicine: Claude Opus 4.7 (86) · Law: Claude Opus 4.7 (84) · Business law: GPT-5 (83).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260607T181746-23cbb6
Started
2026-06-07 18:17 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
6
Judges
GPT-5, Claude Opus 4.8, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.2.3
Pipeline cost
72.24 USD
Judge agreement
68 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7019 / 308178–83+2591,48—
2GPT-5053 / 1467774–81+2720,57—
3Claude Opus 4.8018 / 267471–77—720,93—
4Gemini 2.5 Pro018 / 245450–57+2793,46—
5Mistral Large 2010 / 685046–550494,32—
6Grok 4.304 / 64743–500735,50—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 2,3 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Claude Opus 4.8 at avg 1,4 per answer. Relative to answer length: 2,10 per 1000 output tokens, lowest GPT-5 at 0,71.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77771–82+10,42,32,7100 %2,1078 %
2Claude Opus 4.86963–76—0,60,91,493 %1,0473 %
3GPT-56758–75+60,91,01,9100 %0,7170 %
4Gemini 2.5 Pro3526–44+21,30,51,780 %7,6062 %
5Grok 4.32619–33-71,50,72,1100 %9,4454 %
6Mistral Large 22517–32-12,11,03,0100 %6,3449 %

Medicine · 20 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 95 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,6 per answer. Relative to answer length: 1,17 per 1000 output tokens, lowest Claude Opus 4.8 at 0,57.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78680–91-10,40,81,195 %1,1780 %
2GPT-58275–88+30,40,50,9100 %0,6675 %
3Claude Opus 4.88074–87—0,30,40,685 %0,5775 %
4Mistral Large 26659–7200,50,71,190 %2,2172 %
5Gemini 2.5 Pro5749–64+10,30,40,665 %2,7271 %
6Grok 4.35346–61-20,50,20,775 %3,4970 %

Law · 35 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 86 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,3 per answer. Relative to answer length: 1,10 per 1000 output tokens, lowest GPT-5 at 0,53.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78480–87+20,30,70,986 %1,1080 %
2Claude Opus 4.87972–84—0,30,50,863 %0,8576 %
3GPT-57468–8000,30,60,980 %0,5375 %
4Gemini 2.5 Pro6258–67+10,20,20,340 %1,3669 %
5Grok 4.34842–5400,60,20,863 %6,0966 %
6Mistral Large 24739–54-11,20,82,083 %6,7164 %

Business law · 30 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. Relative to answer length: 0,49 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58377–88+50,20,91,1100 %0,4978 %
2Claude Opus 4.77670–81+40,32,02,397 %1,6070 %
3Claude Opus 4.86761–72—0,61,11,793 %1,1165 %
4Mistral Large 25750–64+30,71,52,197 %3,3754 %
5Gemini 2.5 Pro5144–57+30,70,51,173 %3,9759 %
6Grok 4.35145–57+40,40,81,283 %4,5060 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).