Three new frontier models, one month later. Same 25 questions.
Abstract. In August, four models answered 25 CPA-adjudicated questions on tax judgment, GAAP interpretation, and audit reasoning. In September, three of the four vendors shipped a successor model. We asked the successors the identical 25 questions under the identical protocol: closed book, maximum reasoning, three fresh runs each, every response logged. One model answered 75 of 75. The August leader lost its perfect score to a single one-letter citation. The one unchanged model reproduced its August result exactly, down to the same eleven misses. And one model could not finish two of the questions in nine minutes of trying.
| Model | Current law | Court holdings | Caught-error set | Detection memos | Confident-wrong | Flips | Aug 2026 |
|---|---|---|---|---|---|---|---|
| Claude claude-fable-5-1 | 30/30 | 15/15 | 15/15 | 15/15 | 0 | 0 | 93.3% claude-fable-5 |
| ChatGPT gpt-6-astra | 30/30 | 15/15 | 13/15 | 15/15 | 2 | 1 | 100.0% gpt-5.6-sol |
| Grok grok-4.7 | 15/30 | 15/15 | 6/8 | 15/15 | 7 | 0 | 69.9% grok-4.6 |
| Gemini gemini-3.1-pro-preview | 6/30 | 15/15 | 6/15 | 15/15 | 32 | 0 | 56.0% gemini-3.1-pro-preview |
All runs September 22, 2026, via each provider's API, no browsing, no tools. Grok 4.7 was released the day before. Its 7 non-answers (two items, three attempts each at a 3-minute limit, repeated on a second pass with the same result) are excluded from its denominator and reported in section 04; counting them as wrong, Grok scores 51/75, or 68.0%. A flip means the substantive answer changed between identical runs. The last column is the same vendor's model in round 1. Reading guide, one line: the Confident-wrong column is still the one that matters.
The market leader cited the wrong subparagraph, twice, and quoted statutory text that does not exist.
Question C3 / retired in round 1 / all responses in the exhibit file
Which subparagraph of IRC Section 642(b)(2) provides the $100 deduction for personal exemption allowed to a complex trust?
KEY: subparagraph (A), the $100 general rule. (B) is the $300 rule for trusts required to distribute all income currently.
Nine minutes, twice, and no answer.
Question C2 / retired this round / all responses in the exhibit file
An individual sells Section 1250 property held more than one year. The sale produces an $80,000 net Section 1231 gain before application of the Section 1231(c) lookback rule, of which $25,000 is unrecaptured Section 1250 gain. The taxpayer has $15,000 of nonrecaptured net Section 1231 losses from the five prior years. There is no 28-percent-rate gain. After applying the lookback rule, what amount remains taxed as unrecaptured Section 1250 gain at the maximum 25 percent rate?
KEY: $10,000. Notice 97-59: the recharacterized $15,000 consumes 28-percent gain first (none), then unrecaptured §1250 gain, $25,000 less $15,000.
The exclusion is $15,000,000 by statute. Two models say "about $7 million."
Question A8 / retired this round / all responses in the exhibit file
State the basic exclusion amount applicable to estates of decedents dying in 2026.
KEY: $15,000,000. P.L. 119-21 (2025) set the amount and starts indexing after 2026.
| Vendor | Round 1 model | Round 2 model | Released | Round 1 | Round 2 |
|---|---|---|---|---|---|
| Anthropic | claude-fable-5 | claude-fable-5-1 | Sep 1 | 93.3% | 100.0% |
| OpenAI | gpt-5.6-sol | gpt-6-astra | Sep 3 | 100.0% | 97.3% |
| xAI | grok-4.6 | grok-4.7 | Sep 21 | 69.9% | 75.0% |
| gemini-3.1-pro-preview | same | no release | 56.0% | 56.0% |
Four questions were written down before the run, in a file published with the results. The answers: Anthropic's update fixed every August miss; all five were refusals, and there were none this time. OpenAI's update is not a downgrade on substance, but it is no longer perfect, and its two misses are the same one-letter citation with a fabricated quotation on the second. xAI's update did not leave the second tier: 51 correct in August, 51 in September, and it now times out where 4.6 answered wrong. Google, unchanged, reproduced August exactly: 42 of 75, the same eleven items, every one of them a matter of law that changed in 2025. That last row is the reproducibility check for the whole method. Three of the eleven were published in full in August with their keys. Publishing them changed nothing.
OCR, document extraction, coding, and retrieval-equipped deployments. Models are good at those, and vendors benchmark them constantly. This report is about the chat-box question a partner actually fears: typed question, confident answer, no warning.
Methodology in 60 seconds
25 items, four buckets: current enacted law (10), disguised Supreme Court holdings (5), an archive of errors we have documented frontier models making in our own production logs (5), and staff memos containing exactly one planted defect (5). Every key adjudicated by a licensed CPA against the named primary authority before any model saw anything. The item files, the prompt, and the key are byte-identical to round 1; the hashes below are the August hashes.
Three runs per model per item, each in a fresh conversation, maximum reasoning, no tools, graded independently, no majority voting, no best-of. Binary grading, no partial credit. A failed API call (including a timeout) is recorded as a failure with its error text and is not scored as a wrong answer; that rule, and the four questions this round was designed to answer, were written into a pre-registration file before the first call. Any model answer that disputed our key went back to the primary source before a verdict was entered.
Our interest, disclosed: we sell a CPA test bank that AI writes and a licensed CPA verifies. If models were useless our bank would be worthless; if models were flawless our verification layer would be pointless. We profit only in the middle. This round we published a perfect score by one vendor and a fabricated statute by another. That is what the middle looks like.
Commitment hashes (SHA-256), unchanged from round 1
Frozen core (15 items):7EAD6FAC8C47D54AE041AA71DDFDAD5DB7A71D17858F8620961F09D8119E3147
Current-law tranche (10 items):629A78030C2B744E1D15B9C5E63A860697A28B0386D87A0A53E212D6E4FAABDD
Answer key with grading rules:7A38224F08212276DF6ED35827956A86E734BE89DA618F3B31B348AC4E65F5DA
Round 2 pre-registration (four questions, failure rule, written 22 Sep 2026 before the first call):BD8DF5B901D9F3790C41762AA0F84539D9B9B55682AA159E69574E6CC35AA07D
Every one of the 300 calls: model, snapshot, run number, timestamp, self-reported confidence, verdict, citation verdict, adjudication note, and the 7 non-answers with their error text. Response text is included for the five retired items (A4, A7, C3 from August; C2, A8 from this round) and withheld for active items.
WAVE2-LOG-PUBLIC.CSV WAVE2-EXHIBITS-FULL.CSVwave2-log-public.csv SHA-256: C73EDD4989B71C856846A3A39C8DDB1DC4DCF343D1070D10C8E493DBE8750E9C
wave2-exhibits-full.csv SHA-256: 5AF45FBC043EE273E5946213929BE2659E03102736C7927B854B492DB5DF8F8B
Run by Nicholas Miller, CPA (Oregon #14907). Next round: within two weeks of the next frontier release, same frozen core, same protocol. Standings across all rounds, and every retired question with its key so you can try them in your own chatbot, are on the series page. Corrections to this report, including ones that embarrass us, are posted here, dated, with the original text preserved. Corrections log: none to date.

