AI-Roundtable Leaderboard
Ranking run · 2026-05-12

Ranking run 2026-05-12

Claude Opus 4.7 leads with 77/100 across all domains. Tax law: Claude Opus 4.7 (68) · Medicine: GPT-5 (86) · Law: Claude Opus 4.7 (78) · Business law: Claude Opus 4.7 (75).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260512T135155-40f852
Started
2026-05-12 13:51 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
24.38 USD
Judge agreement
68 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7017 / 267771–81-1541,45—
2GPT-5046 / 716963–76-3630,67—
3Gemini 2.5 Pro022 / 284740–54-2724,20—
3Mistral Large 2016 / 3724740–55-2453,99—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,6 times per answer and adds avg 2,0 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is GPT-5 at avg 2,1 per answer. Relative to answer length: 1,95 per 1000 output tokens, lowest GPT-5 at 0,75.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.76860–76+10,62,02,6100 %1,9580 %
2GPT-55645–67+61,01,12,1100 %0,7573 %
3Gemini 2.5 Pro2514–35-61,50,72,2100 %9,9264 %
4Mistral Large 21710–2502,01,33,3100 %7,1561 %

Medicine · 9 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, GPT-5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,6 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,3 per answer. Relative to answer length: 0,67 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58681–91+40,30,60,9100 %0,6780 %
2Claude Opus 4.78580–90+30,20,60,889 %0,7979 %
3Mistral Large 26964–75+20,60,61,189 %1,9867 %
4Gemini 2.5 Pro5646–66-10,10,10,333 %1,1471 %

Law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,1 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 75 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,4 per answer. Relative to answer length: 1,01 per 1000 output tokens, lowest GPT-5 at 0,80.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77868–86-70,10,70,875 %1,0174 %
2GPT-56047–69-120,60,81,588 %0,8068 %
3Gemini 2.5 Pro5545–65-50,50,10,450 %1,8170 %
4Mistral Large 24844–52-80,50,81,388 %3,7558 %

Business law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,8 times per answer and adds avg 2,4 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is GPT-5 at avg 1,4 per answer. Relative to answer length: 1,71 per 1000 output tokens, lowest GPT-5 at 0,49.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77565–86-20,82,43,2100 %1,7179 %
2GPT-57564–85-90,50,91,4100 %0,4964 %
3Mistral Large 25239–63-11,12,03,2100 %3,9350 %
4Gemini 2.5 Pro5040–58+51,10,71,6100 %4,2957 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).