AI-Roundtable Leaderboard
Ranking run · 2026-05-12

Ranking run 2026-05-12

Claude Opus 4.7 leads with 77/100 across all domains. Tax law: Claude Opus 4.7 (71) · Medicine: Claude Opus 4.7 (85) · Law: Claude Opus 4.7 (80) · Business law: GPT-5 (77).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260512T081552-d5fb45
Started
2026-05-12 08:15 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
24.58 USD
Judge agreement
70 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7017 / 277772–81+2511,56—
2GPT-5029 / 477365–79+3590,72—
3Gemini 2.5 Pro023 / 284842–54+1714,42—
3Mistral Large 2015 / 844839–56+5454,00—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,6 times per answer and adds avg 2,0 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,9 per answer. Relative to answer length: 2,07 per 1000 output tokens, lowest GPT-5 at 0,80.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77164–79+20,62,02,6100 %2,0782 %
2GPT-56349–76+120,91,52,3100 %0,8077 %
3Gemini 2.5 Pro3220–46+11,30,71,988 %7,7965 %
4Mistral Large 2127–17+12,51,13,5100 %7,7460 %

Medicine · 9 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,6 per answer. Relative to answer length: 0,84 per 1000 output tokens, lowest GPT-5 at 0,76.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78575–93+30,20,70,9100 %0,8484 %
2GPT-58578–92+10,30,71,189 %0,7680 %
3Mistral Large 27064–77+30,21,21,4100 %2,2280 %
4Gemini 2.5 Pro5444–63+10,40,30,667 %2,4370 %

Law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,1 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 88 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,4 per answer. Relative to answer length: 1,12 per 1000 output tokens, lowest GPT-5 at 0,68.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78075–85+10,10,91,088 %1,1287 %
2GPT-56444–77-70,90,51,375 %0,6870 %
3Mistral Large 25648–65+90,40,50,975 %2,9560 %
4Gemini 2.5 Pro5141–58-40,40,10,438 %1,9764 %

Business law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,6 times per answer and adds avg 1,1 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. Relative to answer length: 0,64 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57768–84+30,61,11,8100 %0,6471 %
2Claude Opus 4.77163–81+20,62,83,4100 %1,8960 %
3Gemini 2.5 Pro5338–65+70,90,91,8100 %4,9255 %
4Mistral Large 24939–59+41,11,92,9100 %3,8551 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).