AI-Roundtable Leaderboard
AI-Roundtable Leaderboard As of: 10.09.26 Next test: 01.10.26

Ranking of the currently most reliable AI systems for tax, medicine and law.

Reproducible AI-model evaluation for licensed professionals — refreshed monthly. No single model leads everywhere: the per-question winner keeps changing.

Which AI model is the right one this week for my client case, my diagnosis, my contract? The AI-Roundtable Leaderboard delivers the dependable answer: a monthly published ranking of leading language models on real cases from tax law, medicine, law and business law. Triple-judge scoring, hallucination detection, bootstrap confidence intervals, a fully auditable JSONL history.

Recommendation · run of 10 Sep 2026
Claude Opus 5
88 out of 100 points very strong
↑ +2 points stronger than the previous run (86)
Strongest model in 3 of 4 professional fields (exception: Law → Claude Fable 5.1).
4 independent AI reviewers, solid rating consistency (agreement 74%).
Median answer time 63 s — rank 9 of 14 on speed (fastest: Mistral Large 2, 8 s).
CAUTION! The score per model can be misleading! Even the best model changes from question to question — its performance varies even within the same domain. Even the “oracle” (always the best model per question) only reaches an average of 93/100 across all domains — no single model comes close, and a hard remainder stays unsolved. Why this matters …? →
Status & caveats for the current run Run of 10 Sep 2026
Status: Run of 10 Sep 2026, 126 items, all activated domains. 31 published runs since May 2026; trend statements in the history chart from run #4 onward. Methodology complete, raw JSONL transparently inspectable in the audit trail.
Full coverage: GLM-5.3 answered 82/126 items (44 failures, 24 of them timeouts · answer time median 126 s, p90 409 s); Kimi K3 answered 103/126 items (23 failures, 23 of them timeouts · answer time median 146 s, p90 272 s). Diagnostic diff documented in the audit trail with finish_reason=MAX_TOKENS.

Ranking — cross-section across all domains

All tested models by top100 score as a cross-section across all four domains (tax law, medicine, law, business law). The whisker on each bar is the 95 % bootstrap confidence interval — where two whiskers overlap, the order is not reliable (see methodology). The per-domain breakdown follows below — it shows that the order usually differs by subject. So never trust a single model! The R tag before each bar names the reasoning mode the model ran in — always the vendor default, nothing switched off, nothing dialled up (see methodology).

AI-ROUNDTABLE LEADERBOARD ranking.ai-roundtable.de As of: 2026-09-10 02550751001Claude Opus 5Reasoning mode: on — adaptive, the model picks its thinking depth per request (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R adaptive882Claude Fable 5.1Reasoning mode: on — adaptive, the model picks its thinking depth per request (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R adaptive833Claude Opus 4.7Reasoning mode: off. The model has no reasoning mode or ships with it disabled; it runs here exactly as the vendor ships it.R off794GPT-5.6 SolReasoning mode: on — medium, medium effort (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R medium784Kimi K3Reasoning mode: on — standard, vendor standard without an effort dial (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R standard7823 failures, 23 of them timeouts · answer time median 146 s, p90 272 sn=103/1265GPT-5Reasoning mode: on — medium, medium effort (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R medium765GPT-6 AstraReasoning mode: on — medium, medium effort (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R medium766Claude Opus 4.8Reasoning mode: off. The model has no reasoning mode or ships with it disabled; it runs here exactly as the vendor ships it.R off756GLM-5.3Reasoning mode: on — standard, vendor standard without an effort dial (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R standard7544 failures, 24 of them timeouts · answer time median 126 s, p90 409 sn=82/1267Grok 4.5Reasoning mode: on — high, high effort (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R high658Grok 4.6Reasoning mode: on — high, high effort (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R high639DeepSeek V4 ProReasoning mode: on — standard, vendor standard without an effort dial (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R standard5910Gemini 3.1 ProReasoning mode: on — adaptive, the model picks its thinking depth per request (vendor default). Reasoning models get more compute per answer than models without; the ranking compares products as shipped, not architectures.R adaptive5411Mistral Large 2Reasoning mode: off. The model has no reasoning mode or ships with it disabled; it runs here exactly as the vendor ships it.R off48bar = top100 score · whisker = 95 % bootstrap CI · faded = incomplete coverage · shared rank only on identical scoreRanks 1–2 statistically tied (95 % CI overlaps rank 1)R = reasoning mode as shipped by the vendor: off · low · medium · high · adaptive · standard

Score history across all runs

Quality score, rank and vendor comparison per model across the published runs. Retired models stay visible in grey, new entrants start with a marker.

Score history

loading from worker …
Click a legend entry to highlight a model (second click clears, double-click hides). Hovering the chart reads every value at the nearest run.
Table view of all values

Four domains, 126 real case items

Every run scores all activated models on a fixed, versioned question bank (Tax law 31 · Medicine 30 · Law 35 · Business law 30). The items are currently fully synthetic and private — never published, so no vendor can fold them into training. Per model: the score (0–100) as the share of correctly supported target facts, and below it the hallucination bar (scale 0–4 per answer): dark = contradictions of stored key facts (demonstrably wrong), light = additional claims the judge panel flags as false or unsupported. Hover or tap to reveal the numbers.

Tax law · 31 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

hidden profit distribution (vGA), tax-group consolidation (Organschaft), § 8c KStG, cross-border aspects — complex client cases
Claude Opus 5R adaptive89
Claude Fable 5.1R adaptive81
GPT-6 AstraR medium74
Claude Opus 4.7R off73
GPT-5.6 SolR medium73
Claude Opus 4.8R off71
GPT-5R medium65
GLM-5.3R standard59
Grok 4.5R high58
Kimi K3R standard58
Grok 4.6R high53
Gemini 3.1 ProR adaptive49
DeepSeek V4 ProR standard48
Mistral Large 2R off27

In Tax law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 3,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 1,0 per answer. Relative to answer length: 0,59 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,26.

Medicine · 30 Items

Ranks 1–7 statistically tied (95 % CI overlaps rank 1)

multimorbidity, drug interactions, atypical symptom presentation
Claude Opus 5R adaptive92
Claude Fable 5.1R adaptive88
GLM-5.3R standard87
Kimi K3R standard87
Claude Opus 4.7R off85
GPT-5R medium84
GPT-5.6 SolR medium83
GPT-6 AstraR medium82
Claude Opus 4.8R off81
Grok 4.5R high70
DeepSeek V4 ProR standard68
Mistral Large 2R off65
Grok 4.6R high64
Gemini 3.1 ProR adaptive60

In Medicine, Claude Opus 5 (top score) contradicts a stored key fact on average 0,2 times per answer and adds avg 1,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 0,4 per answer. Relative to answer length: 0,65 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,24.

Law · 35 Items

Ranks 1–2 statistically tied (95 % CI overlaps rank 1)

competing claims, standard-terms content review, consequences of a defect of form
Claude Fable 5.1R adaptive91
Claude Opus 5R adaptive91
Kimi K3R standard81
Claude Opus 4.8R off80
Claude Opus 4.7R off79
GPT-5R medium77
GPT-5.6 SolR medium77
GPT-6 AstraR medium73
GLM-5.3R standard70
Grok 4.5R high65
Grok 4.6R high64
DeepSeek V4 ProR standard61
Gemini 3.1 ProR adaptive59
Mistral Large 2R off43

In Law, Claude Fable 5.1 (top score) contradicts a stored key fact on average 0,0 times per answer and adds avg 1,7 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 0,6 per answer. Relative to answer length: 0,60 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,11.

Business law · 30 Items

Ranks 1–9 statistically tied (95 % CI overlaps rank 1)

share-transfer restrictions (Vinkulierung), voting bans, de-facto group liability
Claude Opus 5R adaptive80
GPT-5R medium80
Claude Opus 4.7R off78
GPT-5.6 SolR medium78
GPT-6 AstraR medium75
Kimi K3R standard74
Claude Fable 5.1R adaptive73
Grok 4.6R high69
GLM-5.3R standard68
Claude Opus 4.8R off67
Grok 4.5R high67
DeepSeek V4 ProR standard59
Mistral Large 2R off59
Gemini 3.1 ProR adaptive49

In Business law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 3,2 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 1,0 per answer. Relative to answer length: 0,47 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,25.

✓ Live data · auto-refresh from the currently published run · first bar per domain highlighted green, last red · updated with every new monthly run.

No model leads everywhere — why a single number per model is misleading

A domain percentage is an average across many questions. For each question the best model changes — even within the same domain. The oracle ("best per question": picking the strongest model for every single question) is an upper bound that you cannot hit in advance — and even it stays below 100: a hard remainder that no model solves.

92

Tax law

Oracle 92 (best per question) vs. best single model 89 (Claude Opus 5) — no single model reaches the oracle (+3 points), and the oracle itself stays below 100.

96

Medicine

Oracle 96 (best per question) vs. best single model 92 (Claude Opus 5) — no single model reaches the oracle (+4 points), and the oracle itself stays below 100.

95

Law

Oracle 95 (best per question) vs. best single model 91 (Claude Opus 5) — no single model reaches the oracle (+4 points), and the oracle itself stays below 100.

89

Business law

Oracle 89 (best per question) vs. best single model 80 (Claude Opus 5) — no single model reaches the oracle (+9 points), and the oracle itself stays below 100.

The AI-Roundtable solves the model-selection problem: while it does not reach the (theoretical) "best per question" oracle, it reliably beats every individual model you could fix in advance — because no model is consistently the best across all questions and domains. AI-Roundtable thereby eliminates the user's model-selection risk and delivers a result that sits above every fixed solo model you could choose. Methodology → model selection

Basis: full published monthly run, 30–35 items/domain, all domains — explicitly not the small cross-check sample. Run of 10 Sep 2026.

Maximum safety when it really counts.

Anyone relying on AI in a professional setting cannot afford to work with the wrong model — a fabricated citation, a wrong diagnosis, an embarrassing client letter are too expensive. The AI-Roundtable Leaderboard delivers the dependable answer every month: which model is the right one for the next month.

→

Clear actionable recommendation

One unambiguous model recommendation per domain per run — based on score, hallucination rate and confidence interval. No list to wade through, just a concrete recommendation.

⟳

Monthly rhythm

Model versions change rapidly. The leaderboard keeps you automatically up to date — the recommendation from three months ago is often the wrong one today.

✓

Compliance-conscious presentation

Raw data public, methodology git-versioned, audit trail fully traceable. Demonstrable to regulators, auditors and clients — required reading wherever reliability is demanded.

✉

Straight to your inbox

A snapshot of the most important movements lands in your inbox automatically after every monthly run. The subject line shows the top-mover highlight. Cancel monthly, no account, no login.

What sets the leaderboard apart

Triple-judge scoring

Three independent judge models from three different vendor families (Opus 4.7 · GPT-5 · Mistral Large 2) score every answer source-label-blinded. Inter-rater agreement is part of the public audit value.

Hallucination detection

Every fabricated claim — a wrong case number, an invented statute, a non-existent study — is extracted verbatim and recorded in the audit trail.

Bootstrap confidence intervals

1000 resamples on the cell scores per model × domain. We publish the score mean + 95% percentile — no point value without an uncertainty estimate.

Reproducible at T=0

Question bank git-versioned, model config versioned, deterministic sampling parameters. A second run of the same version produces byte-identical answers.

Append-only history

Once published, run results are never modified or deleted. Methodology changes trigger a new question-bank version; old data points remain visible under their original version.

Raw data public

For each run the complete JSONL — every model answer, every judge rationale — is freely downloadable in the audit trail. No auth, no rate limits.

A perfect complement to the AI-Roundtable app

Anyone running the app on their Mac has several top models debate each other and check one another simultaneously. The leaderboard answers the question that comes before that: which models belong in the roundtable, and which one should take the key role over the next month. The two products work hand in hand.

Frequently asked questions

The key contextual questions about the leaderboard — collapsed, click to expand. If you want to go deeper, you will find the full procedure in the methodology.

There are other ranking studies that deliver completely different results. How does that fit together?

Because they measure something else. The big public leaderboards mostly evaluate generic tasks — multiple-choice knowledge (MMLU), chat preference by gut feeling (LMArena), coding tasks, almost all in English. The AI-Roundtable Leaderboard measures something very narrow and concrete: German-language professional cases from tax law, medicine, law and business law, against manually defined target facts, checked source-label-blinded by three judge families.

A model that tops a generic English best-of list can perform differently on a § 8c-KStG client case or a differential diagnosis — which is exactly what our domain breakdowns show. Three questions decide whether two rankings are comparable at all: what is being measured, on which task type and language, and at which date (model versions change rapidly — hence our monthly rhythm). Our answers to these three questions are disclosed and verifiable in the audit trail; that is the real difference.

On top of that, manual scientific studies often face the problem that the period between conducting the study and publishing the (usually sobering) results can be very long — sometimes more than 1.5 years. That means the study publications you may have just read about in the news largely refer to AI models that were state of the art more than 12 months ago. In today's IT epochs that is literally light-years. But above all, that leaves the user still none the wiser in this very second about what the best model would be RIGHT NOW. That is precisely why our leaderboard ranking measures all the popular top models every month. Only that way do we create real, timely transparency and, above all, trend curves with practical value.

As things stand, AI systems are still worse than human assessments, e.g. on medical questions. Is that true?

For demanding cases: often yes — and that is exactly how we position the leaderboard. The professional remains, for now, the benchmark and the final authority. Our own figures say the same: even the oracle (picking the best model for every question in advance — not achievable in practice) stays below 100 points across all domains. There is a hard remainder that currently no model solves (why a number per model is misleading).

Here AI is a tool for support and a second opinion, not a substitute for professional judgement. The leaderboard helps you choose the most reliable available tool for the next two weeks — substantive responsibility, the counter-review and the final decision remain with the licensed professional. The ranking values are methodically derived guidance, expressly not legal, tax or medical advice.

That said, a further everyday honesty must necessarily be acknowledged: how likely it is that one will reach, with a personal question, a human expert who can actually give "the" desired correct and well-founded answer is equally uncertain. Because not every person is in fact an unrestricted expert in their field. Example: our Roundtable AI model would have passed the Swiss state law examination in every run with a final grade of 1.x on the first attempt. How many people manage that too? In our experience, few. And even fewer are definitively better than an AI — for now. How likely it is that you will land with your matter at an actual expert, the reader may better judge for themselves.

What do I actually do with the results of these rankings in practice?

In practice, in five steps:

  • Choose a model per domain. For the next task, use the model currently ranked highest for your domain (tax / medicine / law / business law) — not "the best model" in general.
  • Don't trust a single model. On important questions, have several models check each other; the gap to the oracle shows that every fixed solo model has gaps.
  • Treat the output as a draft. Every answer is preparatory work; the professional counter-review remains mandatory — especially with a high score, which can feel deceptively safe.
  • Read the hallucination rate alongside it. It calibrates how skeptically you should check citations, case numbers and figures.
  • Check again every month. The recommendation from three months ago is often the wrong one today — model versions shift the order.
How can the reliability of AI systems on important questions be increased?

There are several effective levers that can be combined:

  • Have several independent models cross-check each other. The principle behind triple-judge and the AI-Roundtable app: a finding supported by several models from different families is more dependable than the statement of a single model.
  • Bind answers to sources. Instead of free generation, have the model work with real documents and citations (retrieval, document lookup, web/URL fetch) and counter-check every cited reference.
  • Human in the loop. Final professional review by the licensed professional remains the most important safety anchor.
  • Demand sources and verify them. Explicitly ask for the case number, statute, guideline — and check whether they really exist.
  • Choose the right model for the domain — according to the current ranking, not out of habit.
  • Pseudonymize sensitive data before a model ever sees it (safeguarded in the app via the client mode, Mandatsmodus).
To what extent can hallucinations, for example, be detected or even reduced?

Both are possible — a hallucination cannot be ruled out entirely, but it can be detected and significantly reduced.

Detecting:

  • Consensus across several models. A fact claimed by only a single model and not backed by the others is suspect.
  • Source check. Verify whether the cited case number, statute or referenced study really exists — fabricated citations are the most common case.
  • Machine extraction. Our judges pull every fabricated claim (wrong case no., invented statute, non-existent study) verbatim out of the answer and publish the aggregated hallucination rate per run.

Reducing:

  • Bind to real sources (retrieval / tool use): the model reasons over presented documents instead of from memory.
  • Model debate, in which the models pin each other down on unsupported statements — exactly the procedure of the AI-Roundtable app.
  • Demand and verify citations as well as tightly scoped, well-posed prompts instead of open "tell me about …" questions.

More on hallucination capture in the scoring procedure: Methodology → scoring.

On what empirical basis does this ranking stand?

Before the leaderboard went into operation in April 2026, the scoring pipeline was tested in a multi-month validation study across four professional fields. The current run parameters were only frozen after passing this cross-domain validation.

4

domains scientifically validated

Tax law, medicine (BMJ + multimorbidity cases), law (BGH senates quota-weighted) and business law — each domain with its own question-bank version and a reproducible scoring pipeline.

74 %

judge agreement in the current run

Share of key facts on which all 4 judges return the same verdict (run of 10.09.26). During validation (pilot, May 2026) inter-judge stability was ρ ≥ 0.84; the value in regular operation with four judges and longer answers is lower and is published per run. Phase-E sprints with N=25–100 real case items per run each, tested for inter-judge stability, cross-domain generalization and memorization-confound robustness (pre-/post-training-cutoff comparison).

4

judges from 3 vendor families

GPT-5 · Claude Opus 4.8 · Claude Opus 4.7 · Mistral Large 2; source-label-blinded scoring per answer. Two judges are Anthropic models. Self-preference in the current run (score with all judges minus score with foreign judges only): Anthropic: +1.3 … +2.6 · OpenAI: -4.2 … -2.5 · Mistral: +4.4 points. The headline score stays the all-judge mean; per-model detail on the run page. More detail under Methodology → triple-judge.

The methodological findings established during the validation phase (e.g. Mistral's tax weakness as a domain-stratification effect, the expert-roles lift in BFH, inter-judge stability after tier-1 prompts) feed into the methodology of the running leaderboard and are traceably documented in the full methodology document.

The complete cross-domain validation with all raw statistics, win/loss/tie tables per sub-stratum, limitations and critique response is publicly available as a 10-file gist:

→ Public Validation Gist · 10 files · ~80 sessions cross-domain · as of 2026-05-20

Contains, among others: methodology, legal study (BGH/BFH), medical study (BMJ + MedExpQA), raw-metrics JSON, production roadmap, limitations & threats to validity, critique response, layman abstract. Straight to the limitations file →

Why don't you make the question bank public?
Protection against training contamination. The client items (currently 100 % synthetic, 126 in total) reside exclusively in a private repository and are never published. If the item prompts went public, model vendors could include them in their next training dataset — the ranking would tip from a measurement of genuine generalization ability to a measurement of memorization, and all subsequent runs would be contaminated. For the same reason, in the published raw.jsonl the model answer texts as well as the verbatim-extracted hallucinations are removed — the original case constellations would be reconstructable from them. The aggregated validation results (in the gist above) nonetheless remain fully public — not a single item text in the 10 files, only methodology, sub-stratum statistics and the cross-domain comparison. Complete raw answers are available exclusively on direct NDA request for external auditing.

Recommend it now

Do you know someone who works with AI in a professional setting? The leaderboard is intended as an industry standard for reproducible model evaluation — spreading the word is part of the impact.

Audit trail

Every run is published as a complete, immutable dataset.

Run of 10 Sep 2026 (06:41 UTC)
126 items · 14 models · 4-judge aggregation · inter-rater 0.74
Run ID: run-20260910T064154-bada0b
Live API: https://licenses.ai-roundtable.de/benchmarks/run/run-20260910T064154-bada0b
→ Full run history