AI-Roundtable Leaderboard
Ranking run · 2026-06-20

Ranking run 2026-06-20

Claude Opus 4.7 leads with 80/100 across all domains. Tax law: Claude Opus 4.7 (78) · Medicine: GPT-5.5 Pro (87) · Law: Claude Opus 4.7 (82) · Business law: GPT-5.5 Pro (78).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260620T073912-d2125c
Started
2026-06-20 07:39 UTC
Items
100 (Tax law 15, Medicine 20, Law 35, Business law 30)
Models
8
Judges
GPT-5, Claude Opus 4.8, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-17-v1
Code version
0.2.6
Pipeline cost
122.61 USD
Judge agreement
68 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7020 / 298077–83+2591,48—
2GPT-5.5 Pro n=99/1001199 / 8317875–82-2780,10—
3GPT-5028 / 447672–79+4700,60—
4Claude Opus 4.8018 / 277471–770730,90—
5Kimi K2.6 n=99/1001122 / 2665954–64+2510,25—
6Gemini 3.1 Pro023 / 325450–580754,15—
7Mistral Large 2010 / 685146–56-1504,30—
8Grok 4.305 / 74843–520735,19—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 15 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 1,6 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 1,1 per answer. Relative to answer length: 1,65 per 1000 output tokens, lowest GPT-5.5 Pro at 0,11.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77873–84+70,51,62,1100 %1,6580 %
2GPT-5.5 Pro6860–76+10,90,71,6100 %0,1168 %
3GPT-56754–79+121,11,02,0100 %0,7974 %
4Claude Opus 4.86559–72-20,61,01,6100 %1,1574 %
5Gemini 3.1 Pro4436–5100,80,51,173 %4,6656 %
6Kimi K2.63018–42-31,92,03,9100 %0,4046 %
7Grok 4.32821–37+11,10,71,993 %8,5554 %
8Mistral Large 22718–37+21,91,63,5100 %7,2152 %

Medicine · 20 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Medicine, GPT-5.5 Pro (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 0,3 claims the judge panel flags as false or unsupported; 80 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 0,4 per answer. Relative to answer length: 0,19 per 1000 output tokens, lowest Kimi K2.6 at 0,17.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-5.5 Pro8783–91+10,50,30,780 %0,1981 %
2Claude Opus 4.78375–9000,40,91,3100 %1,2688 %
3Claude Opus 4.88175–87-10,30,40,690 %0,5970 %
4GPT-57870–86-20,60,51,0100 %0,7370 %
5Kimi K2.67063–77+60,40,60,990 %0,1767 %
6Mistral Large 26963–75-20,50,51,085 %1,8972 %
7Gemini 3.1 Pro6357–70+30,30,20,460 %2,0070 %
8Grok 4.36050–68+40,20,40,675 %2,8678 %

Law · 35 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 83 % of answers carry at least one flag. The lowest value is GPT-5.5 Pro at avg 0,6 per answer. Relative to answer length: 1,22 per 1000 output tokens, lowest GPT-5.5 Pro at 0,08.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78277–86+30,30,71,083 %1,2282 %
2Claude Opus 4.88074–85+40,20,50,771 %0,7076 %
3GPT-57773–81+70,30,50,777 %0,3971 %
4GPT-5.5 Pro7770–8300,20,40,671 %0,0878 %
5Kimi K2.66357–69+20,60,51,177 %0,1566 %
6Gemini 3.1 Pro5851–64-10,50,30,760 %3,9771 %
7Grok 4.34943–54+10,50,20,660 %4,9767 %
8Mistral Large 24537–52-41,20,71,880 %6,1062 %

Business law · 30 Items

Ranks 1–4 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5.5 Pro (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,7 claims the judge panel flags as false or unsupported; 97 % of answers carry at least one flag. Relative to answer length: 0,10 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-5.5 Pro7871–84-40,30,71,097 %0,1082 %
2Claude Opus 4.77772–81+10,51,82,3100 %1,7069 %
3GPT-57670–81-30,60,91,5100 %0,6265 %
4Claude Opus 4.86762–73-40,61,01,6100 %1,0863 %
5Kimi K2.66155–67+20,81,92,7100 %0,3152 %
6Mistral Large 25750–63+10,71,32,193 %3,4153 %
7Gemini 3.1 Pro4941–5601,00,61,683 %5,0856 %
8Grok 4.34841–56-30,70,81,587 %5,0957 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).