Medicine: Claude Opus 5 leads with 92/100
14 models on 30 real Medicine cases, scored by 4 independent judge models against stored key facts. Run of 10 Sep 2026. Overlapping confidence intervals mean the order is not reliable.
Ranks 1–7 statistically tied (95 % CI overlaps rank 1)
In Medicine, Claude Opus 5 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 1,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 0,4 per answer. Relative to answer length: 0,65 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,24.
| # | Model | top100 | 95 % CI | Δ prev. | contradicts key fact | extra false/unsupp. | per 1000 tok. |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 92 | 87–96 | +2 | 0,2 | 1,9 | 0,65 |
| 2 | Claude Fable 5.1 | 88 | 83–92 | -2 | 0,2 | 1,3 | 0,71 |
| 3 | GLM-5.3 | 87 | 82–91 | +3 | 0,2 | 1,4 | 0,35 |
| 4 | Kimi K3 | 87 | 81–91 | 0 | 0,2 | 1,3 | 0,50 |
| 5 | Claude Opus 4.7 | 85 | 80–89 | 0 | 0,1 | 1,1 | 1,25 |
| 6 | GPT-5 | 84 | 78–90 | -1 | 0,3 | 1,0 | 0,41 |
| 7 | GPT-5.6 Sol | 83 | 79–88 | 0 | 0,2 | 0,6 | 0,78 |
| 8 | GPT-6 Astra | 82 | 78–85 | -2 | 0,3 | 0,7 | 1,14 |
| 9 | Claude Opus 4.8 | 81 | 75–86 | 0 | 0,2 | 0,7 | 0,87 |
| 10 | Grok 4.5 | 70 | 65–76 | 0 | 0,1 | 0,3 | 1,59 |
| 11 | DeepSeek V4 Pro | 68 | 60–74 | -2 | 0,2 | 0,3 | 0,24 |
| 12 | Mistral Large 2 | 65 | 60–70 | -3 | 0,3 | 1,2 | 2,85 |
| 13 | Grok 4.6 | 64 | 58–70 | +2 | 0,2 | 0,4 | 2,42 |
| 14 | Gemini 3.1 Pro | 60 | 54–65 | 0 | 0,1 | 0,3 | 1,90 |
All values per answer. “contradicts key fact” = active contradiction of a stored fact (demonstrably wrong); “extra false/unsupp.” = claims outside the rubric flagged by the judge panel (union over judges; unsupported counts). Definition: methodology. Full run: run-20260910T064154-bada0b.