<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
<title>AI-Roundtable Ranking — one entry per benchmark run — Tax law</title>
<id>tag:ranking.ai-roundtable.de,2026:/en-feed/steuerrecht.xml</id>
<link rel="self" type="application/atom+xml" href="https://ranking.ai-roundtable.de/en-feed/steuerrecht.xml"/>
<link rel="alternate" type="text/html" href="https://ranking.ai-roundtable.de/"/>
<updated>2026-09-10T06:41:54.000Z</updated>
<author><name>AI-Roundtable Leaderboard</name><uri>https://ranking.ai-roundtable.de</uri></author>
<entry>
<id>tag:ranking.ai-roundtable.de,2026:run:run-20260910T064154-bada0b:steuerrecht</id>
<title>Tax law 2026-09-10: Claude Opus 5 (89)</title>
<link rel="alternate" type="text/html" href="https://ranking.ai-roundtable.de/en-runs/run-20260910T064154-bada0b"/>
<updated>2026-09-10T06:41:54.000Z</updated>
<published>2026-09-10T06:41:54.000Z</published>
<content type="html">&lt;p&gt;&lt;strong&gt;Tax law:&lt;/strong&gt; Claude Opus 5 89/100.&lt;/p&gt;&lt;p&gt;In Tax law, Claude Opus 5 (top score) contradicts a stored key fact on average 0,3 times per answer and adds avg 3,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 1,0 per answer. Relative to answer length: 0,59 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,26.&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://ranking.ai-roundtable.de/en-runs/run-20260910T064154-bada0b&quot;&gt;Full run with confidence intervals, diff to the previous run and raw data&lt;/a&gt;&lt;/p&gt;</content>
</entry>
<entry>
<id>tag:ranking.ai-roundtable.de,2026:run:run-20260907T181839-5742df:steuerrecht</id>
<title>Tax law 2026-09-07: Claude Opus 5 (82)</title>
<link rel="alternate" type="text/html" href="https://ranking.ai-roundtable.de/en-runs/run-20260907T181839-5742df"/>
<updated>2026-09-07T18:18:39.000Z</updated>
<published>2026-09-07T18:18:39.000Z</published>
<content type="html">&lt;p&gt;&lt;strong&gt;Tax law:&lt;/strong&gt; Claude Opus 5 82/100.&lt;/p&gt;&lt;p&gt;In Tax law, Claude Opus 5 (top score) contradicts a stored key fact on average 1,0 times per answer and adds avg 2,4 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 1,2 per answer. Relative to answer length: 0,43 per 1000 output tokens, lowest DeepSeek V4 Pro at 0,26.&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://ranking.ai-roundtable.de/en-runs/run-20260907T181839-5742df&quot;&gt;Full run with confidence intervals, diff to the previous run and raw data&lt;/a&gt;&lt;/p&gt;</content>
</entry>
<entry>
<id>tag:ranking.ai-roundtable.de,2026:run:run-20260801T055440-452791:steuerrecht</id>
<title>Tax law 2026-08-01: Claude Opus 4.7 (75)</title>
<link rel="alternate" type="text/html" href="https://ranking.ai-roundtable.de/en-runs/run-20260801T055440-452791"/>
<updated>2026-08-01T05:54:40.000Z</updated>
<published>2026-08-01T05:54:40.000Z</published>
<content type="html">&lt;p&gt;&lt;strong&gt;Tax law:&lt;/strong&gt; Claude Opus 4.7 75/100.&lt;/p&gt;&lt;p&gt;In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 1,6 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 1,3 per answer. Relative to answer length: 1,60 per 1000 output tokens, lowest GPT-5.6 Sol at 0,27.&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://ranking.ai-roundtable.de/en-runs/run-20260801T055440-452791&quot;&gt;Full run with confidence intervals, diff to the previous run and raw data&lt;/a&gt;&lt;/p&gt;</content>
</entry>
<entry>
<id>tag:ranking.ai-roundtable.de,2026:run:run-20260719T120108-c40839:steuerrecht</id>
<title>Tax law 2026-07-19: Claude Opus 4.7 (78)</title>
<link rel="alternate" type="text/html" href="https://ranking.ai-roundtable.de/en-runs/run-20260719T120108-c40839"/>
<updated>2026-07-19T12:01:08.000Z</updated>
<published>2026-07-19T12:01:08.000Z</published>
<content type="html">&lt;p&gt;&lt;strong&gt;Tax law:&lt;/strong&gt; Claude Opus 4.7 78/100.&lt;/p&gt;&lt;p&gt;In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 1,9 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Grok 4.5 at avg 1,2 per answer. Relative to answer length: 1,72 per 1000 output tokens, lowest GPT-5.6 Sol at 0,27.&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://ranking.ai-roundtable.de/en-runs/run-20260719T120108-c40839&quot;&gt;Full run with confidence intervals, diff to the previous run and raw data&lt;/a&gt;&lt;/p&gt;</content>
</entry>
<entry>
<id>tag:ranking.ai-roundtable.de,2026:run:run-20260621T143914-12112d:steuerrecht</id>
<title>Tax law 2026-06-21: Claude Opus 4.7 (74)</title>
<link rel="alternate" type="text/html" href="https://ranking.ai-roundtable.de/en-runs/run-20260621T143914-12112d"/>
<updated>2026-06-21T14:39:14.000Z</updated>
<published>2026-06-21T14:39:14.000Z</published>
<content type="html">&lt;p&gt;&lt;strong&gt;Tax law:&lt;/strong&gt; Claude Opus 4.7 74/100.&lt;/p&gt;&lt;p&gt;In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 1,8 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Gemini 3.1 Pro at avg 0,9 per answer. Relative to answer length: 1,89 per 1000 output tokens, lowest GPT-5.5 Pro at 0,13.&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://ranking.ai-roundtable.de/en-runs/run-20260621T143914-12112d&quot;&gt;Full run with confidence intervals, diff to the previous run and raw data&lt;/a&gt;&lt;/p&gt;</content>
</entry>
<entry>
<id>tag:ranking.ai-roundtable.de,2026:run:run-20260529T043500-b4ffe0:steuerrecht</id>
<title>Tax law 2026-05-29: Claude Opus 4.7 (71)</title>
<link rel="alternate" type="text/html" href="https://ranking.ai-roundtable.de/en-runs/run-20260529T043500-b4ffe0"/>
<updated>2026-05-29T04:35:00.000Z</updated>
<published>2026-05-29T04:35:00.000Z</published>
<content type="html">&lt;p&gt;&lt;strong&gt;Tax law:&lt;/strong&gt; Claude Opus 4.7 71/100.&lt;/p&gt;&lt;p&gt;In Tax law, Claude Opus 4.7 (top score) contradicts a stored key fact on average 0,5 times per answer and adds avg 2,0 claims the judge panel flags as false or unsupported; 100 % of answers carry at least one flag. The lowest value is Claude Opus 4.8 at avg 1,4 per answer. Relative to answer length: 1,95 per 1000 output tokens, lowest GPT-5 at 0,83.&lt;/p&gt;&lt;p&gt;&lt;a href=&quot;https://ranking.ai-roundtable.de/en-runs/run-20260529T043500-b4ffe0&quot;&gt;Full run with confidence intervals, diff to the previous run and raw data&lt;/a&gt;&lt;/p&gt;</content>
</entry>
</feed>
