Tax law: Claude Opus 5 leads with 89/100
14 models on 31 real Tax law cases, scored by 4 independent judge models against stored key facts. Run of 10 Sep 2026. Overlapping confidence intervals mean the order is not reliable.
Ranks 1–2 statistically tied (95 % CI overlaps rank 1)
In Tax law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 3,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 1,0 per answer. Relative to answer length: 0,59 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,26.
| # | Model | top100 | 95 % CI | Δ prev. | contradicts key fact | extra false/unsupp. | per 1000 tok. |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 89 | 85–92 | +7 | 0,3 | 3,9 | 0,59 |
| 2 | Claude Fable 5.1 | 81 | 74–87 | +6 | 0,4 | 3,6 | 0,67 |
| 3 | GPT-6 Astra | 74 | 68–79 | +5 | 0,3 | 1,6 | 0,98 |
| 4 | Claude Opus 4.7 | 73 | 69–77 | 0 | 0,4 | 2,7 | 2,31 |
| 5 | GPT-5.6 Sol | 73 | 68–78 | +5 | 0,3 | 1,5 | 0,52 |
| 6 | Claude Opus 4.8 | 71 | 66–76 | +2 | 0,3 | 1,7 | 1,38 |
| 7 | GPT-5 | 65 | 57–73 | +3 | 0,8 | 2,0 | 0,45 |
| 8 | GLM-5.3 n=9 | 59 | 42–71 | +13 | 1,3 | 3,8 | 0,54 |
| 9 | Grok 4.5 | 58 | 53–63 | +3 | 0,4 | 0,9 | 3,61 |
| 10 | Kimi K3 n=15 | 58 | 43–71 | -6 | 1,2 | 3,1 | 0,49 |
| 11 | Grok 4.6 | 53 | 46–58 | -7 | 0,6 | 1,3 | 4,72 |
| 12 | Gemini 3.1 Pro | 49 | 43–54 | +4 | 0,4 | 0,6 | 4,22 |
| 13 | DeepSeek V4 Pro | 48 | 41–54 | -2 | 0,8 | 1,5 | 0,26 |
| 14 | Mistral Large 2 | 27 | 21–33 | -2 | 1,4 | 1,9 | 7,22 |
All values per answer. “contradicts key fact” = active contradiction of a stored fact (demonstrably wrong); “extra false/unsupp.” = claims outside the rubric flagged by the judge panel (union over judges; unsupported counts). Definition: methodology. Full run: run-20260910T064154-bada0b.