AI-Roundtable Leaderboard
Ranking run · 2026-05-10

Ranking run 2026-05-10

Claude Opus 4.7 leads with 77/100 across all domains. Tax law: Claude Opus 4.7 (66) · Medicine: Claude Opus 4.7 (86) · Law: Claude Opus 4.7 (78) · Business law: Claude Opus 4.7 (78).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260510T121001-e3a863
Started
2026-05-10 12:10 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
17.13 USD
Judge agreement
75 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1GPT-5 n=16/331717 / 247872–84—770,65—
2Claude Opus 4.7020 / 347772–82—491,61—
3Mistral Large 2015 / 794636–55—414,40—
4Gemini 2.5 Pro n=9/332416 / 333219–47—980,73—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 8 Items

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 1,4 times per answer and adds avg 1,5 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,0 per answer. Relative to answer length: 2,16 per 1000 output tokens, lowest Gemini 2.5 Pro at 0,00.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.76659–73—1,41,52,9100 %2,1677 %
2Mistral Large 293–15—2,51,74,2100 %8,4364 %
3Gemini 2.5 Pro00–0—0,00,00,00 %0,00100 %

Medicine · 9 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 1,0 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,2 per answer. Relative to answer length: 1,43 per 1000 output tokens, lowest GPT-5 at 0,86.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78678–93—0,31,01,4100 %1,4390 %
2GPT-58478–90—0,40,71,189 %0,8679 %
3Mistral Large 26858–77—0,21,01,289 %2,0071 %
4Gemini 2.5 Pro4327–58—0,00,20,250 %1,2878 %

Law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 88 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,0 per answer. Relative to answer length: 1,29 per 1000 output tokens, lowest Gemini 2.5 Pro at 0,00.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77870–85—0,30,81,188 %1,2979 %
2GPT-57063–78—0,10,50,686 %0,4261 %
3Mistral Large 25349–58—0,50,81,388 %3,9858 %
4Gemini 2.5 Pro3811–59—0,00,00,00 %0,0079 %

Business law · 8 Items

In Business law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 2,3 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,0 per answer. Relative to answer length: 1,46 per 1000 output tokens, lowest Gemini 2.5 Pro at 0,00.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77868–88—0,52,32,8100 %1,4670 %
2Mistral Large 25034–65—1,31,62,9100 %4,0356 %
3Gemini 2.5 Pro77–7—0,00,00,00 %0,0090 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).