AI-Roundtable Leaderboard
Ranking run · 2026-05-11

Ranking run 2026-05-11

Claude Opus 4.7 leads with 75/100 across all domains. Tax law: Claude Opus 4.7 (69) · Medicine: GPT-5 (84) · Law: Claude Opus 4.7 (79) · Business law: GPT-5 (74).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260511T051320-09184c
Started
2026-05-11 05:13 UTC
Items
99 (Tax law 24, Medicine 27, Law 24, Business law 24)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
73.04 USD
Judge agreement
67 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7017 / 267572–78-1551,42—
2GPT-5029 / 517066–74-3610,73—
3Gemini 2.5 Pro020 / 244743–51+3684,70—
4Mistral Large 2019 / 884338–48-2404,30—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 24 Items

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,7 times per answer and adds avg 1,7 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 2,0 per answer. Relative to answer length: 1,86 per 1000 output tokens, lowest GPT-5 at 0,97.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.76963–76-20,71,72,4100 %1,8679 %
2GPT-55141–60-41,51,32,896 %0,9770 %
3Gemini 2.5 Pro3124–38-11,40,72,096 %8,6565 %
4Mistral Large 2118–15+22,41,64,0100 %8,2262 %

Medicine · 27 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Medicine, GPT-5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,7 per answer. Relative to answer length: 0,72 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58480–87+20,30,71,0100 %0,7277 %
2Claude Opus 4.78276–86-10,31,01,3100 %1,2379 %
3Mistral Large 26762–70+30,60,71,2100 %2,0566 %
4Gemini 2.5 Pro5348–59-10,40,30,767 %3,0171 %

Law · 24 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 79 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,4 per answer. Relative to answer length: 1,01 per 1000 output tokens, lowest GPT-5 at 0,53.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77973–84+20,30,70,979 %1,0178 %
2GPT-57166–76-20,30,70,975 %0,5372 %
3Gemini 2.5 Pro5549–61+40,30,20,425 %1,8570 %
4Mistral Large 24742–52-70,60,71,392 %4,2560 %

Business law · 24 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 1,1 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. Relative to answer length: 0,60 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57467–80-50,51,11,6100 %0,6063 %
2Claude Opus 4.76965–75-20,71,92,5100 %1,4262 %
3Gemini 2.5 Pro4638–54+81,10,92,092 %5,0657 %
4Mistral Large 24537–54-71,31,93,1100 %3,8151 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).