Can AI be trusted with tax and accounting judgment?
Abstract. Four frontier models answered 25 CPA-adjudicated questions on tax judgment, GAAP interpretation, and audit reasoning. Closed book, each model at its maximum reasoning setting, every question asked three times in fresh conversations, every answer logged: 298 graded responses. One model answered 75 of 75. Another was wrong 44% of the time and attached high confidence to every single miss. The question is no longer whether AI can handle accounting. It is which AI, and whether it warns you when it is guessing.
| Model | Current law | Court holdings | Caught-error set | Detection memos | Confident-wrong | Flips |
|---|---|---|---|---|---|---|
| ChatGPT | 30/30 | 15/15 | 15/15 | 15/15 | 0 | 0 |
| Claude | 30/30 | 14/15 | 15/15 | 11/15 | 0 | 0 |
| Grok | 13/30 | 15/15 | 8/13 | 15/15 | 10 | 2 |
| Gemini | 6/30 | 15/15 | 6/15 | 15/15 | 33 | 2 |
All runs August 22, 2026, via each provider's API, no browsing, no tools (verified in the logged responses). Claude's 5 misses were all refusals to answer, reported separately below; it made zero wrong claims. Grok's 2 timeouts (3 attempts each) are excluded from its denominator and disclosed here. Every model saw the identical prompt. A flip means the substantive answer changed between identical runs; hatched cells above are exactly the flips counted here. Claude answering on one run and refusing on another is reported as refusal inconsistency below, not as a flip. Reading guide, one line: the Confident-wrong column is the one that matters.
The law changed 13 months ago. Two models never noticed.
Question A7 / retired / all 12 responses in the exhibit file
Does Section 199A continue to apply to taxable years beginning in 2026, and what percentage is used in its general deduction formula?
KEY: Yes; 20%. P.L. 119-21 (2025) repealed the scheduled sunset.
Right on the first ask, wrong on the second and third.
Question A4 / retired / all 12 responses in the exhibit file
For a taxpayer who is not married filing separately, state the 2026 SALT deduction limitation, the MAGI threshold at which it begins to decrease, and the minimum limitation after the phase-down.
KEY: $40,400; $505,000; $10,000 floor.
A one-letter citation error, repeated with total confidence.
Question C3 / retired / all 12 responses in the exhibit file
Which subparagraph of IRC Section 642(b)(2) provides the $100 deduction for personal exemption allowed to a complex trust?
KEY: subparagraph (A), the general rule. (B) is the $300 rule for trusts required to distribute all income currently.
OCR, document extraction, coding, and retrieval-equipped deployments. Models are good at those, and vendors benchmark them constantly. This report is about the chat-box question a partner actually fears: typed question, confident answer, no warning.
Methodology in 60 seconds, and why you can trust a test you cannot see
25 items, four buckets: current enacted law (10), disguised Supreme Court holdings (5), an archive of errors we have documented frontier models making in our own production logs (5), and staff memos containing exactly one planted defect (5). Every key adjudicated by a licensed CPA against the named primary authority before any model saw anything.
The questions stay private because published questions leak into training data and the benchmark dies. Instead, the item set was hash-committed before the run: the SHA-256 fingerprints below prove the set existed, unchanged, before any model answered. Retired items are published in full above, with their keys and every model's verbatim answers. Three runs per model per item, each in a fresh conversation, graded independently, no majority voting, no best-of. A refusal grades as incorrect and is reported separately. Binary grading, no partial credit. Any model answer that disputed our key went back to the primary source before a verdict was entered.
Our interest, disclosed: we sell a CPA test bank that AI writes and a licensed CPA verifies. If models were useless our bank would be worthless; if models were flawless our verification layer would be pointless. We profit only in the middle, so our incentives cut both ways. This wave we published a perfect score by the market leader. That is what the middle looks like.
A perfect score at n=25 means no errors detected, not no errors possible. Wave 2 will be harder. That is a promise, not a threat.
Commitment hashes (SHA-256)
Frozen core (15 items, carried to future waves):
7EAD6FAC8C47D54AE041AA71DDFDAD5DB7A71D17858F8620961F09D8119E3147
Wave 1 rotating tranche (10 current-law items):
629A78030C2B744E1D15B9C5E63A860697A28B0386D87A0A53E212D6E4FAABDD
Answer key with grading rules:
7A38224F08212276DF6ED35827956A86E734BE89DA618F3B31B348AC4E65F5DA
Sealed full response log (text released as items retire):
B0F80BB6DE1A67F23AAD9CA281C1FD74C52A8CC13FE9A73B6EC3651A71A7EBF1
Every one of the 298 responses: model, snapshot, run number, timestamp, self-reported confidence, verdict, citation verdict, and adjudication notes. Response text is included for the three retired items and withheld for active items; the sealed log's hash above commits it for future release.
WAVE1-LOG-PUBLIC.CSV WAVE1-EXHIBITS-FULL.CSVwave1-log-public.csv SHA-256:
B2B9AB10480D238CDD891B41ECE6388BCA4FEF3689EDCCE1E8F0AAF9F8614434
wave1-exhibits-full.csv SHA-256:
0A39F02EAFFD3F122956D4D0AE893144B92C05FD199204D40A6A1DE0C5E10D66
Run by Nicholas Miller, CPA (Oregon #14907), who also builds a CPA test bank that AI writes and a CPA verifies, a business that fails if models are useless and fails if models are flawless. The methodology above explains how we mitigate what remains. Next wave: within two weeks of the next frontier model release, on the same frozen core plus a rebuilt current-law tranche. Corrections to this report, including ones that embarrass us, are posted on this page, dated, with the original text preserved. Corrections log: none to date.

