ROUND 2 OF AN ONGOING SERIESALL ROUNDS & STANDINGS  ·  ROUND 1, AUG 2026
THE ACCOUNTING ACCURACY REPORT ROUND 2 / SEP 2026 / N=300 / SAME 25 QUESTIONS AS ROUND 1

Three new frontier models, one month later. Same 25 questions.

Abstract. In August, four models answered 25 CPA-adjudicated questions on tax judgment, GAAP interpretation, and audit reasoning. In September, three of the four vendors shipped a successor model. We asked the successors the identical 25 questions under the identical protocol: closed book, maximum reasoning, three fresh runs each, every response logged. One model answered 75 of 75. The August leader lost its perfect score to a single one-letter citation. The one unchanged model reproduced its August result exactly, down to the same eleven misses. And one model could not finish two of the questions in nine minutes of trying.

17.7%
of 293 answers wrong
79%
of wrong answers carried high confidence
1.0%
flip rate across identical runs
44pt
spread, best to worst model
01Results by model
Claudeclaude-fable-5-1
100.0%
ChatGPTgpt-6-astra
97.3%
Grokgrok-4.7
75.0%
Geminigemini-3.1-pro-preview
56.0%
CORRECT 3/3 MIXED WRONG 3/3 NO ANSWER IN 3 MIN HOVER ANY CELL
02Detail by bucket
ModelCurrent lawCourt holdingsCaught-error setDetection memosConfident-wrongFlipsAug 2026
Claude
claude-fable-5-1
30/3015/1515/1515/150093.3%
claude-fable-5
ChatGPT
gpt-6-astra
30/3015/1513/1515/1521100.0%
gpt-5.6-sol
Grok
grok-4.7
15/3015/156/815/157069.9%
grok-4.6
Gemini
gemini-3.1-pro-preview
6/3015/156/1515/1532056.0%
gemini-3.1-pro-preview

All runs September 22, 2026, via each provider's API, no browsing, no tools. Grok 4.7 was released the day before. Its 7 non-answers (two items, three attempts each at a 3-minute limit, repeated on a second pass with the same result) are excluded from its denominator and reported in section 04; counting them as wrong, Grok scores 51/75, or 68.0%. A flip means the substantive answer changed between identical runs. The last column is the same vendor's model in round 1. Reading guide, one line: the Confident-wrong column is still the one that matters.

03Exhibits: three retired questions, published in full

The market leader cited the wrong subparagraph, twice, and quoted statutory text that does not exist.

Question C3 / retired in round 1 / all responses in the exhibit file

Which subparagraph of IRC Section 642(b)(2) provides the $100 deduction for personal exemption allowed to a complex trust?

KEY: subparagraph (A), the $100 general rule. (B) is the $300 rule for trusts required to distribute all income currently.

CHATGPT RUN 1
Wrong, high confidence: "Subparagraph (B), IRC § 642(b)(2)(B), provides the $100 personal-exemption deduction for a complex trust."
CHATGPT RUN 2
Wrong again, and now with a quotation: "Subparagraph (B) states, ‘Except as provided in subparagraph (C), all other trusts shall be allowed a deduction of $100.’" No subparagraph of § 642(b)(2) says that. The sentence was manufactured to support the wrong letter.
CHATGPT RUN 3
Correct. Same question, same settings, third run: (A).
CLAUDE 3/3
Correct, all runs.
GROK 3/3
Correct, all runs.
GEMINI 0/3
Wrong, all runs, high confidence, (A) and (B) reversed, exactly as in August.
WHY THIS ITEM
This is the one-letter error our own two-model citation audit made in five live simulations earlier this year. In August, Gemini made it three times. In September, the model that scored 100% in August made it twice, and on the second try invented the statute to match. A fabricated quotation is worse than a wrong letter: it survives a reviewer who checks only that the citation is quoted.

Nine minutes, twice, and no answer.

Question C2 / retired this round / all responses in the exhibit file

An individual sells Section 1250 property held more than one year. The sale produces an $80,000 net Section 1231 gain before application of the Section 1231(c) lookback rule, of which $25,000 is unrecaptured Section 1250 gain. The taxpayer has $15,000 of nonrecaptured net Section 1231 losses from the five prior years. There is no 28-percent-rate gain. After applying the lookback rule, what amount remains taxed as unrecaptured Section 1250 gain at the maximum 25 percent rate?

KEY: $10,000. Notice 97-59: the recharacterized $15,000 consumes 28-percent gain first (none), then unrecaptured §1250 gain, $25,000 less $15,000.

CHATGPT 3/3
Correct, all runs. "$10,000 remains unrecaptured Section 1250 gain subject to the maximum 25% rate."
CLAUDE 3/3
Correct, all runs, with the full ordering written out.
GEMINI 3/3
Correct, all runs.
GROK 0/3
No answer. Three attempts of three minutes each, then the whole sequence again on a second pass the same afternoon: six attempts, no response delivered. Same result on C5 (ASC 842 derecognition) and on one of three runs of C1. The other three models answered all three in under a minute every time. We record a non-answer as a failure, not a wrong answer: a practitioner who gets nothing knows to look it up. But nine minutes is longer than looking it up.

The exclusion is $15,000,000 by statute. Two models say "about $7 million."

Question A8 / retired this round / all responses in the exhibit file

State the basic exclusion amount applicable to estates of decedents dying in 2026.

KEY: $15,000,000. P.L. 119-21 (2025) set the amount and starts indexing after 2026.

CHATGPT 3/3
Correct, all runs. "$15,000,000 per decedent."
CLAUDE 3/3
Correct, all runs, noting that indexing begins after 2026.
GEMINI 0/3
Wrong, all runs, high confidence: "reverts to $5 million, adjusted for inflation... broadly estimated to be approximately $7 million." That is the pre-2025 sunset, repealed fourteen months before this test.
GROK 0/3
Wrong, all runs: "Approximately $7 million (inflation-adjusted $5 million statutory base after the temporary increase expires)." Grok rated this Medium confidence. Right instinct, wrong law.
The timeout finding. Grok 4.7 delivered no answer on 7 of 75 calls: all three runs of C2, all three of C5, and one run of C1, the three most computational items in the caught-error set. Each call was allowed three minutes and three attempts, and the entire sequence was repeated on a second pass with the same result. On the 68 calls it did finish, Grok reasoned up to 18,600 tokens on a single question. The pre-registration, written before the run, records a failed call as a failure and not as a wrong answer, so Grok's 75.0% is on answered calls; on all 75 it is 68.0%. One point in Grok's favor: on 10 of its 17 wrong answers it rated its own confidence Medium or Low, and on two current-law items it said outright that it did not have the 2026 figure. That is graded incorrect, because it is not an answer, but it is the behavior the other three should copy.
04What changed since August, and what did not
VendorRound 1 modelRound 2 modelReleasedRound 1Round 2
Anthropicclaude-fable-5claude-fable-5-1Sep 193.3%100.0%
OpenAIgpt-5.6-solgpt-6-astraSep 3100.0%97.3%
xAIgrok-4.6grok-4.7Sep 2169.9%75.0%
Googlegemini-3.1-pro-previewsameno release56.0%56.0%

Four questions were written down before the run, in a file published with the results. The answers: Anthropic's update fixed every August miss; all five were refusals, and there were none this time. OpenAI's update is not a downgrade on substance, but it is no longer perfect, and its two misses are the same one-letter citation with a fabricated quotation on the second. xAI's update did not leave the second tier: 51 correct in August, 51 in September, and it now times out where 4.6 answered wrong. Google, unchanged, reproduced August exactly: 42 of 75, the same eleven items, every one of them a matter of law that changed in 2025. That last row is the reproducibility check for the whole method. Three of the eleven were published in full in August with their keys. Publishing them changed nothing.

05What we did not test

OCR, document extraction, coding, and retrieval-equipped deployments. Models are good at those, and vendors benchmark them constantly. This report is about the chat-box question a partner actually fears: typed question, confident answer, no warning.

06Audit trail
Methodology in 60 seconds

25 items, four buckets: current enacted law (10), disguised Supreme Court holdings (5), an archive of errors we have documented frontier models making in our own production logs (5), and staff memos containing exactly one planted defect (5). Every key adjudicated by a licensed CPA against the named primary authority before any model saw anything. The item files, the prompt, and the key are byte-identical to round 1; the hashes below are the August hashes.

Three runs per model per item, each in a fresh conversation, maximum reasoning, no tools, graded independently, no majority voting, no best-of. Binary grading, no partial credit. A failed API call (including a timeout) is recorded as a failure with its error text and is not scored as a wrong answer; that rule, and the four questions this round was designed to answer, were written into a pre-registration file before the first call. Any model answer that disputed our key went back to the primary source before a verdict was entered.

Our interest, disclosed: we sell a CPA test bank that AI writes and a licensed CPA verifies. If models were useless our bank would be worthless; if models were flawless our verification layer would be pointless. We profit only in the middle. This round we published a perfect score by one vendor and a fabricated statute by another. That is what the middle looks like.

Commitment hashes (SHA-256), unchanged from round 1

Frozen core (15 items):
7EAD6FAC8C47D54AE041AA71DDFDAD5DB7A71D17858F8620961F09D8119E3147

Current-law tranche (10 items):
629A78030C2B744E1D15B9C5E63A860697A28B0386D87A0A53E212D6E4FAABDD

Answer key with grading rules:
7A38224F08212276DF6ED35827956A86E734BE89DA618F3B31B348AC4E65F5DA

Round 2 pre-registration (four questions, failure rule, written 22 Sep 2026 before the first call):
BD8DF5B901D9F3790C41762AA0F84539D9B9B55682AA159E69574E6CC35AA07D

07Download the log

Every one of the 300 calls: model, snapshot, run number, timestamp, self-reported confidence, verdict, citation verdict, adjudication note, and the 7 non-answers with their error text. Response text is included for the five retired items (A4, A7, C3 from August; C2, A8 from this round) and withheld for active items.

WAVE2-LOG-PUBLIC.CSV WAVE2-EXHIBITS-FULL.CSV

wave2-log-public.csv SHA-256: C73EDD4989B71C856846A3A39C8DDB1DC4DCF343D1070D10C8E493DBE8750E9C
wave2-exhibits-full.csv SHA-256: 5AF45FBC043EE273E5946213929BE2659E03102736C7927B854B492DB5DF8F8B

Run by Nicholas Miller, CPA (Oregon #14907). Next round: within two weeks of the next frontier release, same frozen core, same protocol. Standings across all rounds, and every retired question with its key so you can try them in your own chatbot, are on the series page. Corrections to this report, including ones that embarrass us, are posted here, dated, with the original text preserved. Corrections log: none to date.

N. MILLER, CPA / OREGON #14907CHATCPA.IO/AI-ACCURACY