AI-Roundtable Leaderboard
Business law · Run of 2026-09-10

Business law: Claude Opus 5 leads with 80/100

14 models on 30 real Business law cases, scored by 4 independent judge models against stored key facts. Run of 10 Sep 2026. Overlapping confidence intervals mean the order is not reliable.

Ranks 1–9 statistically tied (95 % CI overlaps rank 1)

In Business law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 3,2 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 1,0 per answer. Relative to answer length: 0,47 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,25.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.per 1000 tok.
1Claude Opus 58074–86-10,53,20,47
2GPT-58072–86-10,31,50,33
3Claude Opus 4.77873–82-10,22,82,06
4GPT-5.6 Sol7872–84+20,31,10,56
5GPT-6 Astra7569–82-10,31,20,93
6Kimi K3 n=247465–82+10,53,10,50
7Claude Fable 5.17365–8100,63,20,81
8Grok 4.66962–75+50,41,02,77
9GLM-5.3 n=116855–78-110,63,60,35
10Claude Opus 4.86762–74+10,51,91,56
11Grok 4.56761–71-10,20,82,28
12DeepSeek V4 Pro5952–67+10,51,40,25
13Mistral Large 25952–66+40,42,13,93
14Gemini 3.1 Pro4942–56-20,70,95,07

All values per answer. “contradicts key fact” = active contradiction of a stored fact (demonstrably wrong); “extra false/unsupp.” = claims outside the rubric flagged by the judge panel (union over judges; unsupported counts). Definition: methodology. Full run: run-20260910T064154-bada0b.

← All four domains on the homepage · All runs · Feed (Atom)