AI-Roundtable Leaderboard
Ranking run · 2026-05-10

Ranking run 2026-05-10

Claude Opus 4.7 leads with 78/100 across all domains. Tax law: GPT-5 (90) · Medicine: Claude Opus 4.7 (86) · Law: Claude Opus 4.7 (80) · Business law: Claude Opus 4.7 (74).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260510T164305-2c7b16
Started
2026-05-10 16:43 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
16.65 USD
Judge agreement
74 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7019 / 307874–82+1541,47—
2GPT-5 n=16/331729 / 377669–83-2810,48—
3Mistral Large 2018 / 3704538–54-1384,30—
4Gemini 2.5 Pro n=5/332833 / 351510–23-17961,61—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 1 Items

In Tax law, GPT-5 (top score) contradicts a stored key fact on average 0,0 times per answer and adds avg 0,0 claims the judge panel flags as false or unsupported; 0 % of answers carry at least one flag. Relative to answer length: 0,00 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-59090–90—0,00,00,00 %0,0071 %
2Claude Opus 4.77164–80+50,52,02,5100 %1,8987 %
3Gemini 2.5 Pro1213–13+120,00,30,3100 %3,92100 %
4Mistral Large 2128–15+32,51,33,8100 %8,3657 %

Medicine · 9 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,1 per answer. Relative to answer length: 1,10 per 1000 output tokens, lowest GPT-5 at 0,70.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78681–9200,30,81,1100 %1,1079 %
2GPT-58276–88-20,10,91,078 %0,7068 %
3Mistral Large 26655–76-20,71,01,7100 %2,7067 %
4Gemini 2.5 Pro1913–29-240,00,10,133 %1,0592 %

Law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 88 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,0 per answer. Relative to answer length: 1,23 per 1000 output tokens, lowest Gemini 2.5 Pro at 0,00.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78072–88+20,40,71,188 %1,2380 %
2GPT-56555–77-50,00,50,567 %0,2958 %
3Mistral Large 24941–57-40,80,41,275 %3,5256 %
4Gemini 2.5 Pro77–7-310,00,00,00 %0,0089 %

Business law · 8 Items

In Business law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 2,2 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. Relative to answer length: 1,52 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77465–83-40,52,22,7100 %1,5274 %
2Mistral Large 25341–62+31,41,83,2100 %3,7459 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).