AI-Roundtable Leaderboard
Ranking run · 2026-09-10

Ranking run 2026-09-10

Claude Opus 5 leads with 88/100 across all domains. Tax law: Claude Opus 5 (89) · Medicine: Claude Opus 5 (92) · Law: Claude Fable 5.1 (91) · Business law: Claude Opus 5 (80).

Status
official run (feed, archive, homepage)
Run
run-20260910T064154-bada0b
Started
2026-09-10 06:41 UTC
Items
126 (Tax law 31, Medicine 30, Law 35, Business law 30)
Models
14
Judges
GPT-5, Claude Opus 4.8, Claude Opus 4.7, Mistral Large 2
Question bank
2026-09-09-v5
Code version
0.2.13
Pipeline cost
283.50 USD
Judge agreement
74 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 5063 / 1298886–90+2250,56+2,6
2Claude Fable 5.1049 / 918380–86+1320,70+2,5
3Claude Opus 4.7019 / 297976–81-1461,84+1,3
4GPT-5.6 Sol056 / 1157875–80+2690,57-3,5
4Kimi K3 n=103/12623 (23 T)146 / 2727874–810360,470,0
5GPT-5067 / 1377672–80-1560,37-2,5
5GPT-6 Astra040 / 767673–79+1641,04-4,2
6Claude Opus 4.8019 / 287572–770631,21+1,9
6GLM-5.3 n=82/12644 (24 T)126 / 4097570–79-4300,360,0
7Grok 4.5030 / 496562–68-1792,590,0
8Grok 4.6065 / 1246359–660713,400,0
9DeepSeek V4 Pro0118 / 2565955–63-1660,210,0
10Gemini 3.1 Pro024 / 365451–57-2783,740,0
11Mistral Large 208 / 134844–52-2425,00+4,4

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 31 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 3,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 1,0 per answer. Relative to answer length: 0,59 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,26.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 58985–92+70,33,94,3100 %0,5986 %
2Claude Fable 5.18174–87+60,43,64,0100 %0,6783 %
3GPT-6 Astra7468–79+50,31,61,9100 %0,9872 %
4Claude Opus 4.77369–7700,42,73,0100 %2,3178 %
5GPT-5.6 Sol7368–78+50,31,51,8100 %0,5278 %
6Claude Opus 4.87166–76+20,31,72,0100 %1,3878 %
7GPT-56557–73+30,82,02,7100 %0,4572 %
8GLM-5.35942–71+131,33,85,1100 %0,5473 %
9Grok 4.55853–63+30,40,91,397 %3,6170 %
10Kimi K35843–71-61,23,14,2100 %0,4972 %
11Grok 4.65346–58-70,61,31,997 %4,7269 %
12Gemini 3.1 Pro4943–54+40,40,61,074 %4,2267 %
13DeepSeek V4 Pro4841–54-20,81,52,394 %0,2665 %
14Mistral Large 22721–33-21,41,93,3100 %7,2253 %

Medicine · 30 Items

Ranks 1–7 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 5 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 1,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 0,4 per answer. Relative to answer length: 0,65 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,24.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 59287–96+20,21,92,1100 %0,6590 %
2Claude Fable 5.18883–92-20,21,31,5100 %0,7184 %
3GLM-5.38782–91+30,21,41,6100 %0,3579 %
4Kimi K38781–9100,21,31,6100 %0,5084 %
5Claude Opus 4.78580–8900,11,11,3100 %1,2585 %
6GPT-58478–90-10,31,01,3100 %0,4183 %
7GPT-5.6 Sol8379–8800,20,60,890 %0,7878 %
8GPT-6 Astra8278–85-20,30,71,090 %1,1478 %
9Claude Opus 4.88175–8600,20,70,997 %0,8779 %
10Grok 4.57065–7600,10,30,577 %1,5970 %
11DeepSeek V4 Pro6860–74-20,20,30,580 %0,2472 %
12Mistral Large 26560–70-30,31,21,5100 %2,8569 %
13Grok 4.66458–70+20,20,40,673 %2,4274 %
14Gemini 3.1 Pro6054–6500,10,30,457 %1,9073 %

Law · 35 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Fable 5.1 (top score) contradicts a stored key fact on average 0,0 times per answer and adds avg 1,7 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 0,6 per answer. Relative to answer length: 0,60 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,11.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Fable 5.19188–93+40,01,71,7100 %0,6087 %
2Claude Opus 59188–94+10,11,92,1100 %0,6086 %
3Kimi K38176–86+30,21,72,088 %0,4079 %
4Claude Opus 4.88076–84-10,00,70,780 %0,8077 %
5Claude Opus 4.77973–85-20,21,11,391 %1,4882 %
6GPT-57770–83+20,21,01,286 %0,2877 %
7GPT-5.6 Sol7772–82+20,10,80,983 %0,5576 %
8GPT-6 Astra7367–79+20,31,01,391 %1,2369 %
9GLM-5.37060–79-90,72,02,7100 %0,3279 %
10Grok 4.56559–70-30,20,40,651 %2,8470 %
11Grok 4.66456–71+10,30,40,751 %3,3672 %
12DeepSeek V4 Pro6154–68+20,30,40,763 %0,1172 %
13Gemini 3.1 Pro5952–64-40,20,30,654 %3,1268 %
14Mistral Large 24335–51-21,01,12,189 %7,1564 %

Business law · 30 Items

Ranks 1–9 statistically tied (95 % CI overlaps rank 1)

In Business law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 3,2 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 1,0 per answer. Relative to answer length: 0,47 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,25.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 58074–86-10,53,23,7100 %0,4779 %
2GPT-58072–86-10,31,51,897 %0,3375 %
3Claude Opus 4.77873–82-10,22,83,0100 %2,0673 %
4GPT-5.6 Sol7872–84+20,31,11,5100 %0,5673 %
5GPT-6 Astra7569–82-10,31,21,697 %0,9370 %
6Kimi K37465–82+10,53,13,6100 %0,5073 %
7Claude Fable 5.17365–8100,63,23,7100 %0,8171 %
8Grok 4.66962–75+50,41,01,497 %2,7766 %
9GLM-5.36855–78-110,63,64,2100 %0,3573 %
10Claude Opus 4.86762–74+10,51,92,4100 %1,5669 %
11Grok 4.56761–71-10,20,81,090 %2,2860 %
12DeepSeek V4 Pro5952–67+10,51,41,997 %0,2563 %
13Mistral Large 25952–66+40,42,12,4100 %3,9356 %
14Gemini 3.1 Pro4942–56-20,70,91,597 %5,0758 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).