AI-Roundtable Leaderboard
Ranking run · 2026-05-17

Ranking run 2026-05-17

Claude Opus 4.7 leads with 79/100 across all domains. Tax law: Claude Opus 4.7 (76) · Medicine: Claude Opus 4.7 (84) · Law: Claude Opus 4.7 (80) · Business law: GPT-5 (78).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260517T131406-f96128
Started
2026-05-17 13:14 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
5
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.1.0
Pipeline cost
76.98 USD
Judge agreement
68 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7019 / 287975–82+3591,50—
2GPT-5028 / 497571–78+4670,67—
3Gemini 2.5 Pro018 / 225350–57+9793,66—
4Mistral Large 2018 / 795146–55+6514,20—
5Grok 4.309 / 154642–50+4764,48—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 2,0 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,7 per answer. Relative to answer length: 1,75 per 1000 output tokens, lowest GPT-5 at 0,79.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77668–84+120,32,02,3100 %1,7572 %
2GPT-55947–71+130,91,22,193 %0,7965 %
3Gemini 2.5 Pro3425–44+51,30,61,787 %8,2762 %
4Mistral Large 22920–37+161,61,43,0100 %6,2853 %
5Grok 4.32617–35+81,30,72,080 %8,6561 %

Medicine · 20 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.3 at avg 0,2 per answer. Relative to answer length: 1,25 per 1000 output tokens, lowest GPT-5 at 0,94.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78477–89-30,30,91,2100 %1,2584 %
2GPT-58173–87-20,60,71,295 %0,9480 %
3Mistral Large 26458–69-10,30,91,295 %2,2968 %
4Gemini 2.5 Pro5647–63+40,20,30,545 %2,3777 %
5Grok 4.35246–59-70,10,20,235 %1,0475 %

Law · 35 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 77 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,4 per answer. Relative to answer length: 1,14 per 1000 output tokens, lowest GPT-5 at 0,57.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78074–86+10,20,81,077 %1,1479 %
2GPT-57669–81+20,20,81,080 %0,5771 %
3Gemini 2.5 Pro6256–65+80,30,10,429 %1,7870 %
4Grok 4.34944–54-10,60,20,760 %4,8266 %
5Mistral Large 24739–5401,10,71,783 %5,6765 %

Business law · 30 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 1,0 claims the judge panel flags as false or unsupported; 97 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,2 per answer. Relative to answer length: 0,58 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57872–8300,31,01,497 %0,5870 %
2Claude Opus 4.77569–80+10,51,92,4100 %1,7572 %
3Mistral Large 25751–64+60,71,52,393 %3,6059 %
4Gemini 2.5 Pro5246–59+120,60,61,273 %4,1163 %
5Grok 4.34740–54+60,60,61,270 %4,2259 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).