AI-Roundtable Leaderboard
Ranking run · 2026-05-10

Ranking run 2026-05-10

Claude Opus 4.7 leads with 76/100 across all domains. Tax law: Claude Opus 4.7 (71) · Medicine: Claude Opus 4.7 (83) · Law: Claude Opus 4.7 (77) · Business law: GPT-5 (79).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260510T190628-59fc01
Started
2026-05-10 19:06 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
24.58 USD
Judge agreement
69 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7017 / 317671–80-1531,48—
2GPT-5029 / 387367–78+6650,62—
3Mistral Large 2033 / 3694536–54-1394,43—
4Gemini 2.5 Pro020 / 244438–50+23724,14—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,6 times per answer and adds avg 1,7 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 1,8 per answer. Relative to answer length: 1,70 per 1000 output tokens, lowest GPT-5 at 0,89.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77161–80+30,61,72,3100 %1,7080 %
2GPT-55544–68+131,41,12,5100 %0,8975 %
3Gemini 2.5 Pro3219–46+321,30,81,8100 %7,7361 %
4Mistral Large 294–13+22,91,34,2100 %9,8758 %

Medicine · 9 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,6 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,5 per answer. Relative to answer length: 0,89 per 1000 output tokens, lowest GPT-5 at 0,60.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78373–90+10,30,60,9100 %0,8979 %
2GPT-58275–88-10,20,70,989 %0,6077 %
3Mistral Large 26452–76-40,31,01,3100 %2,2374 %
4Gemini 2.5 Pro5447–61+250,30,30,567 %2,1669 %

Law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,4 per answer. Relative to answer length: 1,41 per 1000 output tokens, lowest GPT-5 at 0,26.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.77769–85-50,30,91,2100 %1,4183 %
2GPT-57366–80+90,00,50,575 %0,2668 %
3Mistral Large 25447–63+10,40,61,088 %3,0162 %
4Gemini 2.5 Pro5143–61+340,30,20,450 %1,9575 %

Business law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 1,3 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. Relative to answer length: 0,60 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-57973–86+30,41,31,7100 %0,6075 %
2Claude Opus 4.77163–80-30,62,63,2100 %1,7163 %
3Mistral Large 25239–6501,32,13,3100 %3,9557 %
4Gemini 2.5 Pro3827–53+181,40,41,7100 %4,4246 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).