AI-Roundtable Leaderboard
Medicine · Run of 2026-09-10

Medicine: Claude Opus 5 leads with 92/100

14 models on 30 real Medicine cases, scored by 4 independent judge models against stored key facts. Run of 10 Sep 2026. Overlapping confidence intervals mean the order is not reliable.

Ranks 1–7 statistically tied (95 % CI overlaps rank 1)

In Medicine, Claude Opus 5 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 1,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 0,4 per answer. Relative to answer length: 0,65 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,24.

#Modeltop10095 % CIΔ prev.contradicts key factextra false/unsupp.per 1000 tok.
1Claude Opus 59287–96+20,21,90,65
2Claude Fable 5.18883–92-20,21,30,71
3GLM-5.38782–91+30,21,40,35
4Kimi K38781–9100,21,30,50
5Claude Opus 4.78580–8900,11,11,25
6GPT-58478–90-10,31,00,41
7GPT-5.6 Sol8379–8800,20,60,78
8GPT-6 Astra8278–85-20,30,71,14
9Claude Opus 4.88175–8600,20,70,87
10Grok 4.57065–7600,10,31,59
11DeepSeek V4 Pro6860–74-20,20,30,24
12Mistral Large 26560–70-30,31,22,85
13Grok 4.66458–70+20,20,42,42
14Gemini 3.1 Pro6054–6500,10,31,90

All values per answer. “contradicts key fact” = active contradiction of a stored fact (demonstrably wrong); “extra false/unsupp.” = claims outside the rubric flagged by the judge panel (union over judges; unsupported counts). Definition: methodology. Full run: run-20260910T064154-bada0b.

← All four domains on the homepage · All runs · Feed (Atom)