AI-Roundtable Leaderboard
Ranking run · 2026-09-04

Ranking run 2026-09-04

Claude Opus 5 leads with 86/100 across all domains. Tax law: Claude Opus 5 (85) · Medicine: Claude Opus 5 (91) · Law: Claude Opus 5 (91) · Business law: Claude Opus 5 (79).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260904T151155-c623e0
Started
2026-09-04 15:11 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
14
Judges
GPT-5, Claude Opus 4.8, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.2.7
Pipeline cost
233.44 USD
Judge agreement
70 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 5062 / 1308683–89—480,39—
2Claude Fable 5.1 n=98/1002 (1 T)49 / 898177–85—400,61—
3Claude Opus 4.7 n=99/1001 (1 T)20 / 297975–82+1601,41—
4Kimi K3 n=81/10019 (18 T)146 / 2867874–83+2500,35—
5GLM-5.3 n=68/10032139 / 4097771–82—470,26—
6Claude Opus 4.8019 / 307673–79+1681,08—
6GPT-5.6 Sol071 / 1817672–79-1670,27—
7GPT-6 Astra0109 / 2147571–78—630,37—
8GPT-5032 / 507469–78+2660,68—
9Grok 4.6068 / 1156562–69—782,63—
10Grok 4.5031 / 526461–68+1832,08—
11DeepSeek V4 Pro n=98/100299 / 2075955–63—710,18—
12Gemini 3.1 Pro n=96/1004 (4 T)27 / 1015551–58+1744,36—
13Mistral Large 208 / 155045–55-1474,64—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

Ranks 1–6 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,7 times per answer and adds avg 2,3 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 1,2 per answer. Relative to answer length: 0,38 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,31.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 58576–94—0,72,33,0100 %0,3880 %
2Claude Fable 5.17664–86—0,92,83,7100 %0,5675 %
3GLM-5.37667–81—0,72,53,2100 %0,3978 %
4Claude Opus 4.77365–82+20,51,82,293 %1,6665 %
5Claude Opus 4.87061–77+50,71,01,6100 %1,1777 %
6GPT-6 Astra6858–77—1,01,32,3100 %0,4074 %
7GPT-5.6 Sol6153–71-71,31,52,7100 %0,3353 %
8GPT-55943–72+51,50,92,5100 %0,8860 %
9Kimi K35634–76-211,41,63,0100 %0,3571 %
10Grok 4.65447–62—1,10,71,893 %4,5159 %
11Grok 4.55042–60-80,80,51,287 %3,5459 %
12Gemini 3.1 Pro4637–55+61,10,71,786 %6,8660 %
13DeepSeek V4 Pro4129–54—1,81,12,892 %0,3156 %
14Mistral Large 22819–39-22,21,43,593 %7,4455 %

Medicine · 20 Items

Ranks 1–9 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 5 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 1,1 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 0,4 per answer. Relative to answer length: 0,45 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,17.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 59185–96—0,41,11,5100 %0,4589 %
2Claude Fable 5.18577–92—0,50,81,3100 %0,6184 %
3GLM-5.38477–90—0,60,91,3100 %0,2782 %
4GPT-58478–90+60,30,60,895 %0,5982 %
5GPT-6 Astra8481–88—0,50,50,995 %0,4978 %
6Claude Opus 4.78376–8900,40,71,095 %0,9272 %
7Claude Opus 4.88377–88+30,30,70,9100 %0,8078 %
8Kimi K38377–90+20,70,81,390 %0,3774 %
9GPT-5.6 Sol8277–87-30,60,30,785 %0,4469 %
10DeepSeek V4 Pro7063–78—0,20,30,465 %0,1766 %
11Grok 4.66861–75—0,30,20,465 %1,8168 %
12Grok 4.56759–75+10,30,20,460 %1,6371 %
13Mistral Large 26560–70-40,60,81,495 %2,8366 %
14Gemini 3.1 Pro6154–68-10,30,20,450 %1,5673 %

Law · 35 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 1,3 claims the judge panel flags as false or unsupported; 94 % of answers carry at least one flag. The lowest value is Grok 4.6 at avg 0,4 per answer. Relative to answer length: 0,47 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,12.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 59187–94—0,31,31,794 %0,4784 %
2Claude Fable 5.18883–92—0,21,31,594 %0,5086 %
3Claude Opus 4.78073–87-20,20,70,979 %1,0482 %
4Claude Opus 4.88076–8400,30,60,980 %0,9678 %
5Kimi K38072–86+50,51,11,688 %0,2979 %
6GPT-5.6 Sol7670–82+10,30,70,983 %0,2472 %
7GLM-5.37464–82—0,81,22,197 %0,2476 %
8GPT-57264–78+10,70,61,280 %0,7372 %
9GPT-6 Astra7164–77—0,90,61,394 %0,4170 %
10Grok 4.66862–74—0,20,20,443 %1,9272 %
11Grok 4.56761–72+20,30,20,451 %2,1871 %
12DeepSeek V4 Pro6154–68—0,50,30,866 %0,1268 %
13Gemini 3.1 Pro5952–65-20,50,30,768 %3,8168 %
14Mistral Large 24638–54+11,10,81,983 %6,3567 %

Business law · 30 Items

Ranks 1–9 statistically tied (95 % CI overlaps rank 1)

In Business law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,7 times per answer and adds avg 1,8 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 0,8 per answer. Relative to answer length: 0,34 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,18.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 57972–86—0,71,82,5100 %0,3465 %
2Claude Opus 4.77873–83+30,52,02,5100 %1,8072 %
3GPT-5.6 Sol7873–83-10,60,81,4100 %0,2270 %
4Kimi K37868–86+30,62,32,9100 %0,4270 %
5GPT-57772–83-20,50,81,393 %0,5563 %
6GPT-6 Astra7671–82—0,70,81,5100 %0,2967 %
7Claude Fable 5.17466–81—0,92,63,5100 %0,7272 %
8GLM-5.37259–84—0,42,73,1100 %0,2672 %
9Claude Opus 4.86964–7500,61,41,997 %1,2569 %
10Grok 4.56661–72+50,30,50,873 %1,6661 %
11Grok 4.66559–71—0,60,71,293 %2,4765 %
12DeepSeek V4 Pro5749–62—0,70,71,390 %0,1864 %
13Mistral Large 25648–63-20,71,52,297 %3,5956 %
14Gemini 3.1 Pro4942–5501,10,61,6100 %5,2157 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).