AI-Roundtable Leaderboard
Ranking run · 2026-05-18

Ranking run 2026-05-18

Claude Opus 4.7 leads with 79/100 across all domains. Tax law: Claude Opus 4.7 (73) · Medicine: Claude Opus 4.7 (87) · Law: Claude Opus 4.7 (83) · Business law: GPT-5 (77).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260518T043700-b4a380
Started
2026-05-18 04:37 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
5
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.1.0
Pipeline cost
77.02 USD
Judge agreement
69 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7021 / 337976–820581,46—
2GPT-5031 / 537572–790690,62—
3Gemini 2.5 Pro019 / 245450–57+1813,32—
4Mistral Large 2013 / 754945–54-2504,15—
5Grok 4.3011 / 164642–500715,08—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 1,8 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,3 per answer. Relative to answer length: 1,71 per 1000 output tokens, lowest GPT-5 at 0,70.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77368–78-30,41,82,2100 %1,7178 %
2GPT-56051–69+11,00,81,8100 %0,7066 %
3Gemini 2.5 Pro4032–50+61,10,41,387 %5,9159 %
4Mistral Large 22415–31-52,11,53,6100 %7,5453 %
5Grok 4.32214–31-41,60,92,587 %10,4359 %

Medicine · 20 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 1,0 claims the judge panel flags as false or unsupported; 90 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,4 per answer. Relative to answer length: 1,21 per 1000 output tokens, lowest GPT-5 at 0,90.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78781–92+30,41,01,390 %1,2183 %
2GPT-58273–88+10,40,81,295 %0,9080 %
3Mistral Large 26456–7100,40,91,290 %2,3367 %
4Gemini 2.5 Pro5648–6400,20,30,445 %2,0176 %
5Grok 4.35143–58-10,40,40,775 %3,1273 %

Law · 35 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 86 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,3 per answer. Relative to answer length: 1,13 per 1000 output tokens, lowest GPT-5 at 0,46.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78379–87+30,20,81,086 %1,1383 %
2GPT-57672–8100,10,60,886 %0,4672 %
3Gemini 2.5 Pro6357–67+10,20,20,343 %1,6472 %
4Grok 4.35449–58+50,40,30,663 %4,2668 %
5Mistral Large 25043–57+30,90,61,477 %4,6268 %

Business law · 30 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,2 per answer. Relative to answer length: 0,60 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57770–84-10,50,91,4100 %0,6074 %
2Claude Opus 4.77367–79-20,42,02,4100 %1,7175 %
3Mistral Large 25245–59-50,81,52,393 %3,6260 %
4Gemini 2.5 Pro4842–55-40,70,51,277 %4,3059 %
5Grok 4.34538–51-20,60,71,390 %4,3553 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).