AI-Roundtable Leaderboard
Ranking run · 2026-05-10

Ranking run 2026-05-10

Claude Opus 4.7 leads with 77/100 across all domains. Tax law: Claude Opus 4.7 (68) · Medicine: GPT-5 (83) · Law: Claude Opus 4.7 (82) · Business law: GPT-5 (76).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260510T175034-5e294d
Started
2026-05-10 17:50 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
21.23 USD
Judge agreement
75 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7020 / 327773–80-1521,50—
2GPT-5034 / 466759–75-9600,74—
3Mistral Large 2016 / 3684637–55+1384,44—
4Gemini 2.5 Pro n=11/332218 / 342113–30+6971,17—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 8 Items

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,6 times per answer and adds avg 1,8 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,0 per answer. Relative to answer length: 1,82 per 1000 output tokens, lowest Gemini 2.5 Pro at 0,00.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.76863–73-30,61,82,5100 %1,8281 %
2GPT-54226–57-482,01,03,0100 %1,1170 %
3Mistral Large 272–13-52,81,74,5100 %9,9656 %
4Gemini 2.5 Pro00–0-120,00,00,00 %0,00100 %

Medicine · 9 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Medicine, GPT-5 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 0,5 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,2 per answer. Relative to answer length: 0,62 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58373–92+10,40,50,9100 %0,6272 %
2Claude Opus 4.78277–87-40,21,01,2100 %1,1983 %
3Mistral Large 26861–75+20,31,01,3100 %2,2075 %
4Gemini 2.5 Pro2919–40+100,00,20,240 %1,7388 %

Law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 1,1 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,0 per answer. Relative to answer length: 1,47 per 1000 output tokens, lowest Gemini 2.5 Pro at 0,00.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78276–87+20,31,11,4100 %1,4776 %
2GPT-56449–78-10,50,51,075 %0,5270 %
3Mistral Large 25349–56+40,50,91,4100 %3,9763 %
4Gemini 2.5 Pro177–34+100,00,00,00 %0,0078 %

Business law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,6 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,0 per answer. Relative to answer length: 0,59 per 1000 output tokens, lowest Gemini 2.5 Pro at 0,00.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57669–83—0,60,91,5100 %0,5974 %
2Claude Opus 4.77465–8200,62,12,7100 %1,4878 %
3Mistral Large 25242–64-11,11,62,7100 %3,4255 %
4Gemini 2.5 Pro2020–20—0,00,00,00 %0,0080 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).