AI-Roundtable Leaderboard
Ranking run · 2026-07-15

Ranking run 2026-07-15

Claude Opus 4.7 leads with 79/100 across all domains. Tax law: Claude Opus 4.7 (72) · Medicine: GPT-5.6 Sol (85) · Law: Claude Opus 4.7 (81) · Business law: GPT-5 (84).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260715T040307-80bf67
Started
2026-07-15 04:03 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
9
Judges
GPT-5, Claude Opus 4.8, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.2.6
Pipeline cost
117.92 USD
Judge agreement
68 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7020 / 297976–820591,44—
2GPT-5032 / 567774–81+2710,61—
2GPT-5.6 Sol n=96/1004 (4 T)91 / 2307773–810700,26—
3Claude Opus 4.8018 / 297471–77-1700,99—
4Kimi K2.6 n=80/10020 (16 T)191 / 3186056–65+2520,29—
5Grok 4.5018 / 305854–62+2753,65—
6Gemini 3.1 Pro020 / 325449–58-1714,80—
7Mistral Large 209 / 155045–54-1484,48—
8GLM-4.6054 / 974540–480650,51—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,7 times per answer and adds avg 2,0 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 1,3 per answer. Relative to answer length: 2,06 per 1000 output tokens, lowest GPT-5.6 Sol at 0,31.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77265–79-30,72,02,7100 %2,0679 %
2GPT-5.6 Sol7052–8500,91,32,2100 %0,3174 %
3Claude Opus 4.86557–72-20,41,41,8100 %1,3270 %
4GPT-56555–7601,10,81,993 %0,7566 %
5Grok 4.54943–56+20,80,51,387 %4,2957 %
6Gemini 3.1 Pro4436–51-40,90,71,680 %7,1870 %
7Kimi K2.63521–49+21,42,03,4100 %0,4062 %
8Mistral Large 22415–35-42,31,74,0100 %8,1348 %
9GLM-4.62316–31+21,61,22,893 %0,7962 %

Medicine · 20 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Medicine, GPT-5.6 Sol (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 0,5 claims the judge panel flags as false or unsupported; 90 % of answers carry at least one flag. The lowest value is GLM-4.6 at avg 0,3 per answer. Relative to answer length: 0,48 per 1000 output tokens, lowest GLM-4.6 at 0,14.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-5.6 Sol8580–89+10,40,50,890 %0,4880 %
2Claude Opus 4.78477–91-40,50,71,2100 %1,1286 %
3GPT-58475–92+30,40,50,990 %0,6478 %
4Claude Opus 4.88073–86-50,30,40,775 %0,6879 %
5Kimi K2.67163–79+20,40,60,995 %0,1971 %
6Mistral Large 26660–71-10,60,51,085 %2,0063 %
7Grok 4.56456–73-30,40,20,575 %1,9469 %
8Gemini 3.1 Pro6356–69-10,30,30,560 %2,3874 %
9GLM-4.66054–67+30,30,10,360 %0,1477 %

Law · 35 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,6 claims the judge panel flags as false or unsupported; 83 % of answers carry at least one flag. The lowest value is Claude Opus 4.8 at avg 0,7 per answer. Relative to answer length: 1,02 per 1000 output tokens, lowest GPT-5.6 Sol at 0,25.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78176–85+10,30,60,983 %1,0277 %
2Claude Opus 4.88177–85+20,20,50,780 %0,7182 %
3GPT-5.6 Sol7569–8000,50,61,086 %0,2569 %
4GPT-57365–80+20,50,61,186 %0,6571 %
5Gemini 3.1 Pro6052–6700,60,30,860 %4,1963 %
6Grok 4.55850–65+60,70,30,969 %5,1466 %
7Kimi K2.65850–6600,90,71,679 %0,2764 %
8GLM-4.64538–52-10,80,51,277 %0,5365 %
9Mistral Large 24537–54-41,30,61,977 %6,4966 %

Business law · 30 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. Relative to answer length: 0,49 per 1000 output tokens, lowest GPT-5.6 Sol at 0,22.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58481–88+30,20,81,1100 %0,4980 %
2Claude Opus 4.77874–83+50,31,92,397 %1,6274 %
3GPT-5.6 Sol7871–83+10,60,81,393 %0,2271 %
4Claude Opus 4.86660–73-30,71,11,8100 %1,1866 %
5Kimi K2.66460–68+30,61,92,5100 %0,3258 %
6Grok 4.55850–64-10,70,61,383 %3,3562 %
7Mistral Large 25750–64+10,71,42,097 %3,2755 %
8Gemini 3.1 Pro4638–55-11,00,71,790 %5,5254 %
9GLM-4.64439–49-30,80,81,690 %0,5253 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).