AI-Roundtable Leaderboard
Ranking run · 2026-05-28

Ranking run 2026-05-28

Claude Opus 4.7 leads with 79/100 across all domains. Tax law: Claude Opus 4.7 (76) · Medicine: Claude Opus 4.7 (87) · Law: Claude Opus 4.7 (82) · Business law: GPT-5 (78).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260528T202802-6ce096
Started
2026-05-28 20:28 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
5
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.2.3
Pipeline cost
76.86 USD
Judge agreement
69 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7020 / 277977–82+3571,53—
2GPT-5039 / 647571–78+1680,64—
3Gemini 2.5 Pro020 / 255248–560774,03—
4Mistral Large 2010 / 695046–55+1494,24—
5Grok 4.308 / 164743–51-1774,53—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 1,9 claims the judge panel flags as false or unsupported; 93 % of answers carry at least one flag. The lowest value is Grok 4.3 at avg 1,5 per answer. Relative to answer length: 1,94 per 1000 output tokens, lowest GPT-5 at 0,73.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77669–83-30,51,92,493 %1,9481 %
2GPT-56151–72+70,81,11,9100 %0,7363 %
3Gemini 2.5 Pro3324–44+31,20,61,787 %7,6265 %
4Grok 4.33324–41+31,10,51,580 %6,5661 %
5Mistral Large 22617–35-11,91,43,3100 %7,0159 %

Medicine · 20 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 1,1 claims the judge panel flags as false or unsupported; 90 % of answers carry at least one flag. The lowest value is Grok 4.3 at avg 0,4 per answer. Relative to answer length: 1,36 per 1000 output tokens, lowest GPT-5 at 0,81.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78782–91+70,41,11,490 %1,3685 %
2GPT-57971–8600,50,61,190 %0,8177 %
3Mistral Large 26659–73-10,40,91,295 %2,2471 %
4Gemini 2.5 Pro5648–6600,30,30,555 %2,3975 %
5Grok 4.35547–6200,30,20,450 %2,2473 %

Law · 35 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 71 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,5 per answer. Relative to answer length: 0,99 per 1000 output tokens, lowest GPT-5 at 0,50.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78279–86+50,20,70,971 %0,9983 %
2GPT-57468–79-30,20,70,871 %0,5070 %
3Gemini 2.5 Pro6155–66-10,30,20,537 %2,3473 %
4Grok 4.34842–54-20,40,20,549 %4,6268 %
5Mistral Large 24841–56+21,00,61,680 %5,1566 %

Business law · 30 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 1,2 claims the judge panel flags as false or unsupported; 93 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,3 per answer. Relative to answer length: 0,65 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57872–8400,41,21,593 %0,6572 %
2Claude Opus 4.77267–7800,62,02,6100 %1,8368 %
3Mistral Large 25446–61+20,71,72,497 %3,8353 %
4Gemini 2.5 Pro4841–56-20,70,61,377 %4,8062 %
5Grok 4.34741–54-20,60,71,383 %4,6655 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).