Business law: Claude Opus 5 leads with 80/100
14 models on 30 real Business law cases, scored by 4 independent judge models against stored key facts. Run of 10 Sep 2026. Overlapping confidence intervals mean the order is not reliable.
Ranks 1–9 statistically tied (95 % CI overlaps rank 1)
In Business law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 3,2 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 1,0 per answer. Relative to answer length: 0,47 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,25.
| # | Model | top100 | 95 % CI | Δ prev. | contradicts key fact | extra false/unsupp. | per 1000 tok. |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 80 | 74–86 | -1 | 0,5 | 3,2 | 0,47 |
| 2 | GPT-5 | 80 | 72–86 | -1 | 0,3 | 1,5 | 0,33 |
| 3 | Claude Opus 4.7 | 78 | 73–82 | -1 | 0,2 | 2,8 | 2,06 |
| 4 | GPT-5.6 Sol | 78 | 72–84 | +2 | 0,3 | 1,1 | 0,56 |
| 5 | GPT-6 Astra | 75 | 69–82 | -1 | 0,3 | 1,2 | 0,93 |
| 6 | Kimi K3 n=24 | 74 | 65–82 | +1 | 0,5 | 3,1 | 0,50 |
| 7 | Claude Fable 5.1 | 73 | 65–81 | 0 | 0,6 | 3,2 | 0,81 |
| 8 | Grok 4.6 | 69 | 62–75 | +5 | 0,4 | 1,0 | 2,77 |
| 9 | GLM-5.3 n=11 | 68 | 55–78 | -11 | 0,6 | 3,6 | 0,35 |
| 10 | Claude Opus 4.8 | 67 | 62–74 | +1 | 0,5 | 1,9 | 1,56 |
| 11 | Grok 4.5 | 67 | 61–71 | -1 | 0,2 | 0,8 | 2,28 |
| 12 | DeepSeek V4 Pro | 59 | 52–67 | +1 | 0,5 | 1,4 | 0,25 |
| 13 | Mistral Large 2 | 59 | 52–66 | +4 | 0,4 | 2,1 | 3,93 |
| 14 | Gemini 3.1 Pro | 49 | 42–56 | -2 | 0,7 | 0,9 | 5,07 |
All values per answer. “contradicts key fact” = active contradiction of a stored fact (demonstrably wrong); “extra false/unsupp.” = claims outside the rubric flagged by the judge panel (union over judges; unsupported counts). Definition: methodology. Full run: run-20260910T064154-bada0b.