AI-Roundtable Leaderboard
Ranking run · 2026-05-14

Ranking run 2026-05-14

Claude Opus 4.7 leads with 75/100 across all domains. Tax law: Claude Opus 4.7 (68) · Medicine: Claude Opus 4.7 (82) · Law: Claude Opus 4.7 (77) · Business law: GPT-5 (80).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260514T075711-8acb3a
Started
2026-05-14 09:23 UTC
Items
99 (Tax law 24, Medicine 27, Law 24, Business law 24)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
73.57 USD
Judge agreement
67 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7020 / 317572–78-2491,59—
2GPT-5032 / 567066–74+1620,68—
3Mistral Large 2016 / 794641–51-1384,36—
4Gemini 2.5 Pro023 / 274541–48-2724,25—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 24 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,8 times per answer and adds avg 2,1 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,8 per answer. Relative to answer length: 2,09 per 1000 output tokens, lowest GPT-5 at 0,87.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.76862–7400,82,12,8100 %2,0978 %
2GPT-55344–62-31,11,32,496 %0,8772 %
3Gemini 2.5 Pro3023–37+51,40,51,888 %7,8159 %
4Mistral Large 2139–17-42,21,53,7100 %7,6262 %

Medicine · 27 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,4 per answer. Relative to answer length: 1,15 per 1000 output tokens, lowest GPT-5 at 0,69.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78276–87-30,30,91,2100 %1,1583 %
2GPT-58075–84-60,40,61,096 %0,6968 %
3Mistral Large 26762–72-20,60,91,596 %2,4469 %
4Gemini 2.5 Pro5449–57-20,30,20,474 %1,9271 %

Law · 24 Items

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 83 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,5 per answer. Relative to answer length: 1,35 per 1000 output tokens, lowest GPT-5 at 0,54.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77771–82-10,40,81,283 %1,3573 %
2GPT-56661–70+60,30,61,071 %0,5464 %
3Gemini 2.5 Pro5145–57-40,30,20,546 %2,3870 %
4Mistral Large 24945–54+10,50,81,288 %3,7859 %

Business law · 24 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 1,1 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. Relative to answer length: 0,57 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58076–84+50,51,11,6100 %0,5774 %
2Claude Opus 4.77265–78-30,62,32,9100 %1,6271 %
3Mistral Large 25143–58-11,52,03,5100 %4,3258 %
4Gemini 2.5 Pro4437–51-61,20,61,792 %4,6944 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).