AI-Roundtable Leaderboard
Ranking run · 2026-05-12

Ranking run 2026-05-12

Claude Opus 4.7 leads with 78/100 across all domains. Tax law: Claude Opus 4.7 (67) · Medicine: Claude Opus 4.7 (82) · Law: Claude Opus 4.7 (85) · Business law: GPT-5 (84).

Status
development run — not in the feed, not on the homepage (roster test / interim run)
Run
run-20260512T105829-f0ac95
Started
2026-05-12 10:58 UTC
Items
33 (Tax law 8, Medicine 9, Law 8, Business law 8)
Models
4
Judges
GPT-5, Claude Opus 4.7, Mistral Large 2
Question bank
2026-05-08-pilot
Code version
0.1.0
Pipeline cost
24.58 USD
Judge agreement
71 %

Overall across domains

#Modeltop10095 % CIΔ prev.failuressec median/p90hallu bandper 1000 tok.Δ own family
1Claude Opus 4.7017 / 257873–82+1491,59—
2GPT-5031 / 507265–79-1640,66—
3Gemini 2.5 Pro023 / 274942–55+1734,18—
3Mistral Large 2017 / 3664941–57+1423,91—

Δ prev. = difference to the previous published run; “—” = model was not in it. Hallu band: 100 = none, 0 = ≥ 4 false claims per answer. Per 1000 tok. = flagged claims per 1000 output tokens (length-normalised). Δ own family = score with all judges minus score with foreign judges only, in points; positive = own family scores more generously; “—” = not recorded before code_version 0.2.8 or no foreign judge (methodology §4, §6.5).

Tax law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,8 times per answer and adds avg 2,0 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 2,0 per answer. Relative to answer length: 2,01 per 1000 output tokens, lowest GPT-5 at 0,91.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.76757–75-40,82,02,8100 %2,0183 %
2GPT-55035–63-131,31,32,5100 %0,9165 %
3Gemini 2.5 Pro3118–46-11,30,72,088 %8,8171 %
4Mistral Large 2179–25+52,41,53,9100 %8,7554 %

Medicine · 9 Items

Ranks 1–3 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,7 times per answer and adds avg 0,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,5 per answer. Relative to answer length: 1,55 per 1000 output tokens, lowest GPT-5 at 0,76.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78273–88-30,70,91,6100 %1,5582 %
2GPT-58275–89-30,30,71,0100 %0,7676 %
3Mistral Large 26759–74-30,40,91,3100 %2,0075 %
4Gemini 2.5 Pro5745–68+30,20,30,567 %2,0463 %

Law · 8 Items

In Law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,0 times per answer and adds avg 0,8 claims the judge panel flags as false or unsupported; 75 % of answers carry at least one flag. The lowest value is Gemini 2.5 Pro at avg 0,1 per answer. Relative to answer length: 0,98 per 1000 output tokens, lowest GPT-5 at 0,35.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1Claude Opus 4.78581–90+50,00,80,875 %0,9889 %
2GPT-57263–81+80,10,50,675 %0,3572 %
3Gemini 2.5 Pro6053–68+90,10,00,125 %0,3872 %
4Mistral Large 25650–6200,30,50,775 %2,0959 %

Business law · 8 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Business law, GPT-5 (top score) contradicts a stored key fact on average 0,4 times per answer and adds avg 1,3 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. Relative to answer length: 0,57 per 1000 output tokens.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.avg totalanswers affectedper 1000 tok.IRR
1GPT-58479–90+70,41,31,6100 %0,5781 %
2Claude Opus 4.77771–84+60,62,32,9100 %1,5976 %
3Mistral Large 25338–64+41,42,13,5100 %3,8056 %
4Gemini 2.5 Pro4532–58-81,40,51,8100 %5,2360 %

Raw data of this run

raw.jsonl and progress.jsonl are the public streams with response texts stripped (contamination protection, methodology §7).