AI-Roundtable Leaderboard
Ranking run · 2026-05-16

Ranking run 2026-05-16

Claude Opus 4.7 leads with 76/100 across all domains. Tax law: Claude Opus 4.7 (64) · Medicine: Claude Opus 4.7 (87) · Law: Claude Opus 4.7 (79) · Business law: GPT-5 (78).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260516T200347-bc6991
Started
2026-05-16 20:03 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
5
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
28.68 USD
Judge agreement
68 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7022 / 367671–82-2551,40—
2GPT-5049 / 667162–78-2610,71—
3Mistral Large 2023 / 3664537–53+1404,33—
4Gemini 2.5 Pro019 / 224438–51-2734,23—
5Grok 4.307 / 104234–51+1605,81—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,8 times per answer and adds avg 1,5 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,8 per answer. Relative to answer length: 1,73 per 1000 output tokens, lowest GPT-5 at 0,87.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.76458–72-50,81,52,3100 %1,7377 %
2GPT-54626–65-151,60,82,4100 %0,8765 %
3Gemini 2.5 Pro2916–43+31,40,51,888 %8,3664 %
4Grok 4.3184–3502,41,03,3100 %11,6856 %
5Mistral Large 2138–1902,01,73,7100 %7,7759 %

Medicine · 9 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 89 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,6 per answer. Relative to answer length: 0,93 per 1000 output tokens, lowest GPT-5 at 0,71.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78779–94+60,20,81,089 %0,9388 %
2GPT-58377–89-20,20,81,0100 %0,7185 %
3Mistral Large 26558–71-30,31,01,3100 %2,2370 %
4Grok 4.35949–69+70,30,40,744 %2,9473 %
5Gemini 2.5 Pro5241–6200,30,40,667 %2,7268 %

Law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,0 times per answer and adds avg 1,0 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,2 per answer. Relative to answer length: 1,14 per 1000 output tokens, lowest GPT-5 at 0,47.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77968–88-60,01,01,0100 %1,1476 %
2GPT-57463–84+60,10,70,863 %0,4777 %
3Gemini 2.5 Pro5444–63+10,10,10,225 %0,9975 %
4Grok 4.35042–57+20,30,30,463 %2,8065 %
5Mistral Large 24738–58-21,00,61,588 %4,9463 %

Business law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 1,5 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,7 per answer. Relative to answer length: 0,70 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57872–85-10,51,52,0100 %0,7071 %
2Claude Opus 4.77463–84-30,52,53,0100 %1,5964 %
3Mistral Large 25140–63+81,61,73,2100 %3,8254 %
4Grok 4.34124–58-31,30,82,0100 %4,7155 %
5Gemini 2.5 Pro4029–51-101,30,51,7100 %4,7051 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).