The same 25 questions. Every frontier release.
Whenever OpenAI, Anthropic, xAI, or Google ships a new frontier model, we retest all four on the same 25 CPA questions on tax, GAAP, and audit. Each model gets three fresh attempts per question, 75 per round, and a licensed CPA grades every returned answer. In September, Claude scored 75 of 75. GPT scored 73 of 75; both misses cited the wrong subparagraph, and the second fabricated a quote from the tax code: 2002 wording plus a clause that never existed, presented as current law. Grok returned 68 of 75 attempts and got 51 right: 75% of returned answers, 68% of attempts. Gemini, unchanged since August, missed the same 11 questions, 33 wrong answers, exactly as before.
One page · PDFThe results on a single pageStandings, the key finding, the three public questions with keys. For a class or a partner meeting.DOWNLOAD ↓| Vendor | All rounds | Sep 2026 → | Aug 2026 → | Confidently wrong all rounds |
|---|---|---|---|---|
| OpenAIChatGPT | 98.7% | 97.3%−2.7 · gpt-6-astra | 100.0%gpt-5.6-sol | 1.3% |
| AnthropicClaude | 96.7% | 100.0%+6.7 · claude-fable-5-1 | 93.3%claude-fable-5 | 0.0% |
| xAIGrok | 72.3% | 75.0%+5.1 · 51/68 · grok-4.7 | 69.9%grok-4.6 | 12.1% |
| GoogleGemini | 56.0% | 56.0%±0 · gemini-3.1-pro-preview | 56.0%gemini-3.1-pro-preview | 43.3% |
Accuracy is correct answers divided by returned answers, out of 75 attempts per round. Refusals count as wrong, because the model responded. A call that returns no answer (a timeout or API error) is not graded: it is left out of the figure and reported separately: Grok returned 68 of 75 in September, so its September figure is 51 of 68 and its all-rounds figure is 102 of 141. Its +5.1 is the smaller denominator, not more right answers: 51 correct in both rounds.
All rounds follows the vendor across model versions, not one model over time; the model that sat each month is named in the cell. Confidently wrong is the share of returned answers that were wrong and labeled High confidence. Rows are ordered by all-rounds accuracy. We sell a CPA-verified test bank; method, hashes, and conflict of interest are in section 06.
One email when a new round is published: the newest models, the same 25 questions. Nothing else.
A current enacted law, 2026 figures. B Supreme Court holdings with the facts disguised. C known traps: errors we have seen frontier models make in production. D staff memos with exactly one planted defect.
Read down a column for a model, across a row for a question. A1 to A8 are the law that changed in 2025 (A9 and A10 are GAAP items): two vendors have now missed most of them twice, a month apart, with new models in between. C3 is a one-letter citation three different models have gotten wrong. B and D are nearly solid black: models reason well about facts they are given and fail on facts that changed after they were trained.
Ranked by cost per correct answer, not by accuracy. Same round as the standings. In this round, the lower-priced models were also the less accurate ones.
| Model | Correct / answered | Spend, 75 calls (Grok: 68) | Output tokens per call | Cost per correct answer |
|---|---|---|---|---|
| Grokgrok-4.7 | 51 / 68 | $1.73 | 3,398 | $0.034 |
| Geminigemini-3.1-pro-preview | 42 / 75 | $1.72 | 1,872 | $0.041 |
| Claudeclaude-fable-5-1 | 75 / 75 | $6.96 | 1,791 | $0.093 |
| ChatGPTgpt-6-astra | 73 / 75 | $8.17 | 2,135 | $0.112 |
Token counts from each provider's own API response, priced at list rates on the day of the run (Claude and GPT $10 in / $50 out per million; Gemini $2 / $12; Grok $2 / $6). Reasoning tokens are billed as output by all four and are included. Grok's spend covers only the 68 calls that returned; xAI does not report usage on an aborted call. In this round the gap between the cheapest right answer and the most reliable one was 6 cents. A wrong answer a partner has to unwind costs more than that.
Round 2 · September 22, 2026 · gpt-6-astra, claude-fable-5-1, grok-4.7, gemini-3.1-pro-preview
Three new frontier models, one month later. Same 25 questions.
Claude Fable 5.1 answered 75 of 75. GPT-6 Astra lost its perfect score to a single one-letter citation, and on the second miss fabricated a quote from the tax code: 2002 wording plus a clause that never existed, presented as current law. Grok 4.7 matched 4.6's 51 correct, timed out on seven attempts, and never finished two of the questions. Gemini, unchanged, reproduced August exactly: same score, same eleven misses.
READ ROUND 2 →Round 1 · August 22, 2026 · gpt-5.6-sol, claude-fable-5, grok-4.6, gemini-3.1-pro-preview
GPT-5.6 went 75 for 75. Grok and Gemini sat the 2026 exam with 2024 law.
GPT-5.6 answered 75 of 75. Claude made zero wrong claims but refused five times. Grok and Gemini answered as if the 2025 tax law had never passed, and 72% of all wrong answers carried high confidence. The round that set the protocol: hashes, three runs, binary grading, refusals and timeouts reported on their own.
READ ROUND 1 →Teaching this? Download the one-page instructor brief (PDF, rounds 1 to 2): standings, the key finding, these three questions with keys, and discussion prompts. Free to use in class.
Three of the 25 have been published in full, with keys, and are still asked every round. Paste one into any chatbot and compare. If it gets A7 or A4 wrong, it is answering from before the 2025 law. C3 is a pinpoint citation nobody memorizes; it is here because models misplace or invent citations with confidence. ("Complex trust" in C3 means a trust not required to distribute all income currently; citing (A) and noting the (B) exception for trusts that are grades correct.) Every model's verbatim responses to these are in the exhibit files, linked in section 06.
Does Section 199A continue to apply to taxable years beginning in 2026, and what percentage is used in its general deduction formula?
KEY: Yes; 20%. P.L. 119-21 (2025) repealed the scheduled sunset.
For a taxpayer who is not married filing separately, state the 2026 SALT deduction limitation, the MAGI threshold at which it begins to decrease, and the minimum limitation after the phase-down.
KEY: $40,400; phase-down begins at $505,000 MAGI; floor $10,000.
Which subparagraph of IRC Section 642(b)(2) provides the $100 deduction for personal exemption allowed to a complex trust?
KEY: Subparagraph (A), the $100 general rule. (B) is the $300 rule for trusts required to distribute all income currently.
Method, in 60 seconds
25 items, four buckets: current enacted law (10), disguised Supreme Court holdings (5), known traps from our production logs (5), and staff memos containing exactly one planted defect (5). Every key adjudicated by a licensed CPA against the named primary authority before any model saw anything.
The questions stay private because published questions leak into training data and the benchmark dies. Three (A4, A7, C3) are published on purpose, as contamination controls: they stay in the run, and if models start getting these right while still missing the private ones, that is the leak showing. The item set was hash-committed before the first run; the fingerprints below prove it existed, unchanged, before any model answered. Three runs per model per item, each in a fresh conversation, no tools, no majority voting, no best-of, binary grading. A refusal grades as incorrect and is reported separately. A call that returns no answer (a timeout or API error) is not graded; it is left out of the accuracy figure and reported separately. Each round's questions and rules are written down before the run.
Our interest, disclosed: we sell a CPA test bank that AI writes and a licensed CPA verifies. If models were useless our bank would be worthless; if models were flawless our verification layer would be pointless. We profit only in the middle, so our incentives cut both ways.
The data: every call, every response
Every call in every round: model, snapshot, run, timestamp, answer, confidence and verdict. The exhibit files hold the full verbatim responses to the published questions. Check any file against its SHA-256.
Round 2 (Sep 2026)wave2-log-public.csvSHA-256 457E3D6F58E02A9C9312F70D7DA38EBB06BF50F141037951155449753996CCA7wave2-exhibits-full.csvSHA-256 634C28012A92290F9257271ED578DDFF0FC1191EF08CAA5E3DA23FC83C3881FA
Round 1 (Aug 2026)wave1-log-public.csvSHA-256 B2B9AB10480D238CDD891B41ECE6388BCA4FEF3689EDCCE1E8F0AAF9F8614434wave1-exhibits-full.csvSHA-256 0A39F02EAFFD3F122956D4D0AE893144B92C05FD199204D40A6A1DE0C5E10D66
Commitment hashes (SHA-256)
Frozen core (15 items):7EAD6FAC8C47D54AE041AA71DDFDAD5DB7A71D17858F8620961F09D8119E3147
Current-law tranche (10 items):629A78030C2B744E1D15B9C5E63A860697A28B0386D87A0A53E212D6E4FAABDD
Answer key with grading rules:7A38224F08212276DF6ED35827956A86E734BE89DA618F3B31B348AC4E65F5DA
Run by Nicholas Miller, CPA (Oregon #14907). The next round runs within two weeks of the next frontier release, on the same 25 questions. Get it by email. Corrections, including ones that embarrass us, are posted on the round pages, dated, with the original text preserved.

