AI-Roundtable Leaderboard
Tax law · Run of 2026-09-10

Tax law: Claude Opus 5 leads with 89/100

14 models on 31 real Tax law cases, scored by 4 independent judge models against stored key facts. Run of 10 Sep 2026. Overlapping confidence intervals mean the order is not reliable.

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Tax law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 3,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 1,0 per answer. Relative to answer length: 0,59 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,26.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.per 1000 tok.
1Claude Opus 58985–92+70,33,90,59
2Claude Fable 5.18174–87+60,43,60,67
3GPT-6 Astra7468–79+50,31,60,98
4Claude Opus 4.77369–7700,42,72,31
5GPT-5.6 Sol7368–78+50,31,50,52
6Claude Opus 4.87166–76+20,31,71,38
7GPT-56557–73+30,82,00,45
8GLM-5.3 n=95942–71+131,33,80,54
9Grok 4.55853–63+30,40,93,61
10Kimi K3 n=155843–71-61,23,10,49
11Grok 4.65346–58-70,61,34,72
12Gemini 3.1 Pro4943–54+40,40,64,22
13DeepSeek V4 Pro4841–54-20,81,50,26
14Mistral Large 22721–33-21,41,97,22

All values per answer. “contradicts key fact” = active contradiction of a stored fact (demonstrably wrong); “extra false/unsupp.” = claims outside the rubric flagged by the judge panel (union over judges; unsupported counts). Definition: methodology. Full run: run-20260910T064154-bada0b.

← All four domains on the homepage · All runs · Feed (Atom)