AI-Roundtable Leaderboard
Law · Run of 2026-09-10

Law: Claude Fable 5.1 leads with 91/100

14 models on 35 real Law cases, scored by 4 independent judge models against stored key facts. Run of 10 Sep 2026. Overlapping confidence intervals mean the order is not reliable.

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

In Law, Claude Fable 5.1 (top score) contradicts a stored key fact on average 0,0 times per answer and adds avg 1,7 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 0,6 per answer. Relative to answer length: 0,60 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,11.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.per 1000 tok.
1Claude Fable 5.19188–93+40,01,70,60
2Claude Opus 59188–94+10,11,90,60
3Kimi K3 n=348176–86+30,21,70,40
4Claude Opus 4.88076–84-10,00,70,80
5Claude Opus 4.77973–85-20,21,11,48
6GPT-57770–83+20,21,00,28
7GPT-5.6 Sol7772–82+20,10,80,55
8GPT-6 Astra7367–79+20,31,01,23
9GLM-5.3 n=327060–79-90,72,00,32
10Grok 4.56559–70-30,20,42,84
11Grok 4.66456–71+10,30,43,36
12DeepSeek V4 Pro6154–68+20,30,40,11
13Gemini 3.1 Pro5952–64-40,20,33,12
14Mistral Large 24335–51-21,01,17,15

All values per answer. “contradicts key fact” = active contradiction of a stored fact (demonstrably wrong); “extra false/unsupp.” = claims outside the rubric flagged by the judge panel (union over judges; unsupported counts). Definition: methodology. Full run: run-20260910T064154-bada0b.

← All four domains on the homepage · All runs · Feed (Atom)