AI-Roundtable Leaderboard
Ranking run · 2026-07-11

Ranking run 2026-07-11

Claude Opus 4.7 leads with 79/100 across all domains. Tax law: Claude Opus 4.7 (75) · Medicine: Claude Opus 4.7 (88) · Law: Claude Opus 4.7 (80) · Business law: GPT-5 (81).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260711T103548-436ac1
Started
2026-07-11 10:35 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
9
Judges
GPT-5, Claude Opus 4.8, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.2.6
Pipeline cost
119.54 USD
Judge agreement
68 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7020 / 317976–82-1601,41—
2GPT-5.6 Sol079 / 1917773–80—670,28—
3Claude Opus 4.8019 / 307572–780720,95—
3GPT-5031 / 467571–79-1700,60—
4Kimi K2.6 n=78/10022 (22 T)169 / 2495853–63-1480,27—
5Grok 4.5016 / 315652–61—723,99—
6Gemini 3.1 Pro020 / 285551–59-1744,34—
7Mistral Large 209 / 165147–560534,15—
8GLM-4.6 n=97/1003 (3 T)61 / 864540–49-2650,48—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 1,5 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 1,1 per answer. Relative to answer length: 1,54 per 1000 output tokens, lowest GPT-5.6 Sol at 0,29.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77567–82+20,41,51,9100 %1,5475 %
2GPT-5.6 Sol7061–78—1,11,42,5100 %0,2969 %
3Claude Opus 4.86762–73+10,70,81,6100 %1,1268 %
4GPT-56556–74+51,00,81,887 %0,6573 %
5Gemini 3.1 Pro4841–5500,90,31,173 %4,6863 %
6Grok 4.54738–55—1,10,71,793 %5,1456 %
7Kimi K2.63319–45+21,91,53,4100 %0,3459 %
8Mistral Large 22821–35+11,71,12,8100 %6,1352 %
9GLM-4.62113–2902,10,82,8100 %0,8964 %

Medicine · 20 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 90 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 0,3 per answer. Relative to answer length: 1,04 per 1000 output tokens, lowest GLM-4.6 at 0,13.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78883–92+40,30,81,190 %1,0487 %
2Claude Opus 4.88581–89+50,40,40,685 %0,5980 %
3GPT-5.6 Sol8479–88—0,50,30,780 %0,4277 %
4GPT-58173–88-10,40,61,095 %0,7079 %
5Kimi K2.66957–79-50,70,41,080 %0,2174 %
6Grok 4.56758–76—0,30,10,355 %1,2466 %
7Mistral Large 26760–73+20,50,61,0100 %2,0764 %
8Gemini 3.1 Pro6458–70-10,20,30,460 %1,6770 %
9GLM-4.65752–64-20,10,20,340 %0,1374 %

Law · 35 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 77 % of answers carry at least one flag. The lowest value is Claude Opus 4.8 at avg 0,6 per answer. Relative to answer length: 1,11 per 1000 output tokens, lowest Kimi K2.6 at 0,20.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78075–85-20,20,81,077 %1,1179 %
2Claude Opus 4.87975–82-30,20,40,666 %0,7078 %
3GPT-5.6 Sol7568–82—0,50,61,189 %0,3177 %
4GPT-57163–78-30,60,61,186 %0,6469 %
5Gemini 3.1 Pro6052–67-10,50,20,754 %4,0370 %
6Kimi K2.65851–65-70,80,81,683 %0,2064 %
7Grok 4.55244–60—0,90,31,163 %6,6772 %
8Mistral Large 24940–5700,90,81,777 %5,9068 %
9GLM-4.64639–53-50,70,51,168 %0,4259 %

Business law · 30 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 97 % of answers carry at least one flag. Relative to answer length: 0,50 per 1000 output tokens, lowest GPT-5.6 Sol at 0,22.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58174–86-10,40,81,297 %0,5074 %
2GPT-5.6 Sol7771–83—0,50,91,4100 %0,2268 %
3Claude Opus 4.77368–79-40,61,82,5100 %1,7466 %
4Claude Opus 4.86963–7500,61,21,897 %1,2170 %
5Kimi K2.66155–68+40,82,02,8100 %0,3555 %
6Grok 4.55950–67—0,70,71,380 %3,4459 %
7Mistral Large 25649–6200,81,32,197 %3,5056 %
8Gemini 3.1 Pro4739–56-21,20,71,897 %5,6752 %
9GLM-4.64740–54-10,81,01,7100 %0,5157 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).