Law: Claude Fable 5.1 leads with 91/100
14 models on 35 real Law cases, scored by 4 independent judge models against stored key facts. Run of 10 Sep 2026. Overlapping confidence intervals mean the order is not reliable.
Ranks 1–2 statistically tied (95 % CI overlaps rank 1)
In Law, Claude Fable 5.1 (top score) contradicts a stored key fact on average 0,0 times per answer and adds avg 1,7 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 0,6 per answer. Relative to answer length: 0,60 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,11.
| # | Model | top100 | 95 % CI | Δ prev. | contradicts key fact | extra false/unsupp. | per 1000 tok. |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | 91 | 88–93 | +4 | 0,0 | 1,7 | 0,60 |
| 2 | Claude Opus 5 | 91 | 88–94 | +1 | 0,1 | 1,9 | 0,60 |
| 3 | Kimi K3 n=34 | 81 | 76–86 | +3 | 0,2 | 1,7 | 0,40 |
| 4 | Claude Opus 4.8 | 80 | 76–84 | -1 | 0,0 | 0,7 | 0,80 |
| 5 | Claude Opus 4.7 | 79 | 73–85 | -2 | 0,2 | 1,1 | 1,48 |
| 6 | GPT-5 | 77 | 70–83 | +2 | 0,2 | 1,0 | 0,28 |
| 7 | GPT-5.6 Sol | 77 | 72–82 | +2 | 0,1 | 0,8 | 0,55 |
| 8 | GPT-6 Astra | 73 | 67–79 | +2 | 0,3 | 1,0 | 1,23 |
| 9 | GLM-5.3 n=32 | 70 | 60–79 | -9 | 0,7 | 2,0 | 0,32 |
| 10 | Grok 4.5 | 65 | 59–70 | -3 | 0,2 | 0,4 | 2,84 |
| 11 | Grok 4.6 | 64 | 56–71 | +1 | 0,3 | 0,4 | 3,36 |
| 12 | DeepSeek V4 Pro | 61 | 54–68 | +2 | 0,3 | 0,4 | 0,11 |
| 13 | Gemini 3.1 Pro | 59 | 52–64 | -4 | 0,2 | 0,3 | 3,12 |
| 14 | Mistral Large 2 | 43 | 35–51 | -2 | 1,0 | 1,1 | 7,15 |
All values per answer. “contradicts key fact” = active contradiction of a stored fact (demonstrably wrong); “extra false/unsupp.” = claims outside the rubric flagged by the judge panel (union over judges; unsupported counts). Definition: methodology. Full run: run-20260910T064154-bada0b.