Methodology
How the questions get made, and what gets deleted.
Anyone can ask a model for a hundred questions. Finding its errors is the product. Here is the process, the numbers, and the part it cannot catch.
Every attempt in the three builds behind this bank is published. (One earlier build predates these logs; it shipped nothing and was written off entirely, and its cost is disclosed below.) 27,312 questions generated, 9,883 deleted at review, one pulled after shipping (the Section 179 question, logged in the changelog), and 223 that predate these logs: 17,658 in the bank, and the arithmetic ties to the question. The simulations: 947 build attempts, 344 deleted, 603 shipped. Each row carries its timestamp, topic, verdicts, what it cost, and for the simulations the field-by-field agreement of two blind solvers.
March build, 26,162 rows (CSV, 4.9 MB) August build, 1,150 rows (CSV, 289 KB) Simulation build, 944 rows (CSV, 543 KB) Citation audit, 1,808 rows (CSV, 351 KB)
All four files carry a published SHA-256 further down this page, so you can confirm the file you downloaded is the file we published.
And the questions themselves. A log tells you what was rejected, not whether what survived is any good. So sixty questions are published in full, ten from each section, with every option, the keyed answer and a separate written rationale under each wrong choice. Including one question that was wrong, shipped, and had to be withdrawn. No account, nothing collapsed, nothing behind a script.
| # | Step | Who | Sees the answer | May edit | Produces |
|---|---|---|---|---|---|
| 0 | Design | CPA | — | yes | The Blueprint mapping, the angle to test, and the question architecture the writer must follow |
| 1 | Write | Model A | writes it | yes | One question, four options, four explanations |
| 2 | Judge | Model B | yes | no | Accept, or reject with a reason |
| 3 | Blind solve | Model C | no | no | An answer, plus answerability and ambiguity judgments |
| 4 | Blind solve | Model D | no | no | An answer, plus answerability and ambiguity judgments |
| 5 | Compare | Arithmetic | — | no | Ship, or delete |
| 6 | Sign off, then keep reading | CPA | yes | yes | Correlated-error check, ongoing |
On step 6, because the honest version matters. The timestamps in the log cover generation only, a 16-day window in March. Review is not a stage that finished when the build did. Every one of the 17,658 has been read, post-build, finishing August 12, 2026, by the licensed CPA who signs this page. The daily reading continues as re-review, which is why the changelog keeps getting entries rather than stopping at launch. The precise claim, because precision is the point: every question has been read and signed off. The independent model re-check of each one was done at build time, not by me — no human could re-work 17,658 alone, and I will not claim to have.
Each model is from a different family, and no model ever grades its own work. The table shows the pipeline as it runs today: since the August 2026 builds the solver gate is fully blind — solvers receive the stem and four options only, with the key and the explanations stripped server-side, and return an answer plus judgments on answerability and ambiguity. The March build's second gate worked with open books instead: its two models read the full item, explanations included, and independently accepted or rejected it, which is why that gate's reject codes below include explanation conflicts. On that build the control was independence; blindness was added after it. A rejected question is never repaired, it is deleted and written again from scratch, because editing a flawed question preserves the shape that made it flawed. Since the August build, every question is also commissioned against one specific Blueprint point, not a topic, so questions from the same topic cannot collapse into each other.
| Condition | Required | Fails to |
|---|---|---|
| Blind solver C vs blind solver D | same letter | delete |
| Blind letter vs recorded answer | match | review by hand |
| Answerable as written | yes, both | delete |
| Blocking defect reported | none, both | delete |
| Second defensible answer | none, both | delete |
| Judge verdict | accept | regenerate once |
| Model confidence score | not used | — |
Confidence is deliberately excluded. In this build a model returned 99 out of 100 on a question with two defensible answers. A confidence floor also risked deleting good hard questions, if models express less confidence on the items that discriminate hardest.
| Why 9,741 questions were deleted | Gate 1 | Gate 2 |
|---|---|---|
| Missing a fact needed to decide the answer | 2,761 | 614 |
| Technical error | 2,241 | 1,683 |
| Two defensible answers | 1,082 | 440 |
| Explanation conflicts with the key or the stem | 634 | 445 |
| Correct only under superseded guidance | 486 | 191 |
| Stem contradicts itself | 404 | 719 |
| Misleading wording | 302 | 368 |
| Unsafe date or threshold | 127 | — |
| Other trust risk | 1 | 127 |
| Rejected by this gate | 8,038 | 4,587 |
The two gates fail differently. Gate one's most common finding was a missing decisive fact; gate two's was a technical error. They overlapped on only 2,884 questions.
| The build, 15 to 30 March 2026 | Count | Share |
|---|---|---|
| Questions generated | 26,162 | — |
| Questions deleted | 9,741 | 37.2% |
| Questions kept | 16,421 | 62.8% |
| Second build, 5 to 7 August 2026: filling our own coverage gaps | Count | Share |
|---|---|---|
| Questions generated | 1,150 | — |
| Questions deleted at the gate | 142 | 12.3% |
| Questions shipped | 1,008 | 87.7% |
| Pulled after shipping, on review | 1 | — |
| Net added to the bank | 1,007 | — |
| Cost, per question generated | $0.138 | — |
| Cost, total | $159 | — |
| Cost per question that survived | $0.158 | — |
Both builds together: $3,889 across 27,312 attempts. That is 14.2 cents an attempt and 22.3 cents for each of the 17,428 questions those two builds put in the bank. The August batch cost less per surviving question, 15.8 cents against 22.7, for one reason: fewer of them had to be thrown away.
| Simulation build, 9 to 11 August 2026 | Count | Share |
|---|---|---|
| Build attempts | 944 | — |
| Deleted: failed generation, review discards, final validation | 344 | 36.4% |
| Simulations shipped | 600 | 63.6% |
| Graded fields shipped | 8,395 | — |
| Cost, total | $427 | — |
| Cost, per attempt | $0.45 | — |
| Cost per simulation that shipped | $0.71 | — |
Why a shipped simulation costs 71 cents against a question's 22. A simulation carries about fourteen graded fields, and two blind solvers work every field before it ships, with disagreements sent to review. More surface, more checking. The reject rate landing within a quarter point of the question builds' is not a target we set; it is what the same standard costs wherever it is applied.
| The whole pipeline | Written | Deleted | Reject rate | Cost | Per shipped |
|---|---|---|---|---|---|
| Questions, both builds | 27,312 | 9,883 | 36.2% | $3,889 | $0.22 |
| Simulations | 947 | 344 | 36.3% | $429 | $0.71 |
| Total | 28,259 | 10,227 | 36.2% | $4,318 | $0.24 |
Reperform the bottom row: 28,259 written and 10,227 deleted come from adding the three logs above, the reject rate is one divided by the other, and $4,318 is $3,889 plus the $428.83 that the simulation log's cost column sums to. Per shipped is total cost over the 18,031 items these logs put in the bank.
Why the second build's discard rate is a third of the first. Same gate, same rules, a better generator. Both builds started from the same human design work: a CPA mapping each topic to the Blueprint and specifying the angles to test and the architecture to follow. The difference is granularity. March prompts carried that design at the topic level; after we read what the 9,741 rejections had in common, the prompts were rewritten so that every August question was commissioned against one unique Blueprint point. A 12.3% discard rate is what the same checking finds when the writing has already been fixed once.
One of the 1,008 did not survive review. A Section 179 question computed its phase-out on thresholds that changed in July 2025. Both blind solvers agreed on the wrong answer and the judge passed it, because all three models used the same outdated figure. Unanimity is not validation when the error is correlated, which is the one failure this gate cannot catch by design. It was pulled on 7 August and the entry is in the changelog.
The wrong-answer warranty. A paying customer who is first to show that a keyed answer is wrong under the law date its question states is paid $50, keeps their access, and the mistake is published in the changelog. The rules are in the terms. Everything on this page is why we can afford the offer.
| Where the money went, March build | Total | Per attempt | Share |
|---|---|---|---|
| Writing the questions | $1,425 | $0.054 | 38% |
| Checking them: judge, two independent reviewers, comparison and orchestration | $2,305 | $0.089 | 62% |
| Total | $3,730 | $0.143 | 100% |
| Cost per question that survived | — | $0.227 | — |
| Earlier build, written off entirely | $2,000 | — | — |
Generating the questions was 38 percent of the cost. Checking them was the other 62. That ratio is the entire difference between this and a bank someone produced by asking a model for a hundred questions two hundred times. Five and a half cents makes a question. Twenty-three cents puts one in this bank: the full pipeline cost per survivor, carrying the review of every attempt and the generation cost of the 9,741 that were deleted.
The full log is published. Every one of the 26,162 attempts, with its timestamp, section, topic, both gate verdicts, the reject code, and what it cost. One honest limit: the March log's schema predates per-row model columns, so the model identities for that build are disclosed on this page rather than in the file; the August and simulation logs carry them on every row. Download the CSV (4.8 MB), or read the summary JSON if you just want the totals.
CSV SHA-256:
dca855fd123ac340490d666b834fd5ed800a8f9fb56b02464a8cc9f2e6df553a
Published so you can confirm the file you downloaded is the file we published, and so we cannot
quietly swap it later.
The second build's log is published too. All 1,150 attempts from 5 to 7 August, with both gate verdicts, the reject reason, and the reference of every question that shipped. Download the August CSV, or the summary JSON. It also flags the 98 rows where our own gate-2 output came back as malformed JSON and was initially counted as a rejection; recovering those is why the discard rate fell from 15.7% to 12.3% between our first count and the published one.
August CSV SHA-256:
0e01aae0dacf8befff9aa848875665667a74a7f4fdfa52fc79c0fb5b59eb57df
References are assigned when a question is loaded into the bank, not when it is generated, so each log records what was written and judged rather than what was ultimately loaded. The bank holds 17,658 questions: 16,421 from the March build, 1,007 from the August build, 223 that predate both, and 7 added by hand since.
Added by hand since the logs. Small batches written outside a numbered build, usually to catch a change in the law. They go through the same blind-solve gate. They are listed here rather than given their own build log, because a log with seven rows is paperwork, not evidence.
| Date | Questions | Section | Why |
|---|---|---|---|
| September 13, 2026 | 7 | REG | OBBBA business provisions: bonus depreciation, Section 179, Section 174A, Section 199A |
| Total | 7 | ||
| # | Step | Who | Sees the key | Produces |
|---|---|---|---|---|
| 1 | Write | Anthropic model | writes it | One simulation: scenario, exhibits, 6–16 graded fields, each with its own answer, explanation and authority citation, plus a written method for attacking that kind of simulation |
| 2 | Blind solve | OpenAI model | no | An answer for every field, key stripped server-side |
| 3 | Blind solve | Google model | no | An answer for every field, key stripped server-side |
| 4 | Compare, field by field | Arithmetic | — | Ship only if both solvers reproduce the key; numeric fields within a stated tolerance |
| 5 | Review | CPA | yes | Keep, or kill — kills are in the log with the reason |
A simulation that failed verification was rebuilt from scratch, never repaired. After two failures an item shipped nothing and waited for a rebuilt attempt under an improved prompt; the log shows every one of those retries individually.
| The simulation build, 9 to 11 August 2026 | Count |
|---|---|
| Build attempts | 944 |
| Attempts where the blind solvers reproduced the entire key — 642 with both solvers, 7 with one completed solver, marked clean_single in the log | 649 |
| Discarded on CPA review after passing — early-prompt builds, naming and dating defects, reviewer kills | 30 |
| Duplicate clean builds of the same simulation, best attempt kept | 8 |
| Excluded by final mechanical validation | 4 |
| Discarded as unfit on the final pre-ship read | 7 |
| Simulations shipped | 600 |
| Graded fields shipped by this build | 8,395 |
| Cost, total | $427 |
| Cost per shipped simulation | $0.71 |
Reperform the waterfall: 649 passing attempts minus 30 review discards is 619 rows, which collapse to 611 unique simulations once the 8 double-builds keep their best attempt; final mechanical validation cut 4 and the pre-ship read cut 7, leaving the 600 that shipped. The log carries 11 further discard marks on attempts that had already failed solving, and one attempt was lost to a spreadsheet cell-size cap before reaching any verdict; neither group is inside the 649. That lost attempt also carries no model name in its Writer Model field, so the summary JSON counts 943 attempts by model string and 944 by vendor. The log records what the machines did; the final-validation and pre-ship cuts above are the human layer, documented here rather than back-edited into the file. Every one of the 600 shipped references traces to a passing attempt in the log.
Gap fill, September 2, 2026. A teardown of a competitor's EPS simulation sent us back to our own coverage, where we found the same gap on our side: one simulation on public company reporting topics, none on earnings per share, none on contracts. Three attempts through the same pipeline, three shipped clean on the first try: weighted-average shares and basic EPS, diluted EPS with antidilutive-security screening, and contract formation with statute of frauds and agent authority. 40 graded fields, $1.38, bringing the file total to $428.83. The log carries the three new rows, the hash below is updated to match, and the running totals are 947 attempts and 603 shipped.
The cost ratio inverted. On the multiple-choice build, writing was 38% of the spend and checking was 62%. On simulations it is 89% writing and 11% checking, because a simulation is expensive to author — exhibits, a full answer key, an explanation and citation for every field — and comparatively cheap to blind-solve. The check stayed just as strict; the work moved.
Seven of the 600 shipped on a single solver. Where one solver failed for
infrastructure reasons and the other reproduced the key in full, the row says clean_single
rather than clean. They are marked in the log; judge them as you see fit.
Citation granularity is a policy, not an accident. Every citation is written at the finest level that is stable and independently verifiable. IRC citations carry full subsection precision (§1222(1) vs §1222(3) is the content) and ASC citations carry paragraph precision, because that numbering is stable. Auditing standards are cited at section plus topic — "AU-C 315 (identifying and assessing risks of material misstatement)" — because SAS 143 and SAS 145 rebuilt the paragraph numbering of the revised standards, and paragraph-level claims there are no longer reliably checkable by anyone, including the AI models that review this bank. The blind-solve gate verifies answers, not citations, so the citation layer is checked separately: by the CPA read-through, and by an independent two-model audit whose corrections land in the changelog.
The full log is published. Every attempt with its models, token counts, cost, field-agreement counts and verdict, plus the reviewer's decision column. Dispute details are withheld only on rows whose simulation is live in the paid bank, so the log cannot be used to extract answer keys; rejected attempts are fully readable. Download the simulation CSV (543 KB), or the summary JSON if you just want the totals.
Reperform it. Every number in this tab is recomputable from the CSV: the row count is the attempts, counting the Verdict column reproduces the verdict split, summing the Total Cost column reproduces the spend, counting Decision = discard reproduces the review discards, and the per-topic shipped counts on the blueprint page sum to 603. The summary JSON states these checks explicitly so they can be run by a person or a machine.
Simulation CSV SHA-256:
90c126b51579d3e04ada4feb84a07fe5e229ff35f40e33e6eb518c7dfa038712
The citation audit ledger is published too. Every citation the August 2026 audit changed: 349 corrections (unanimous, adjudicated, and class-sweep, each row labeled with how it was caught) plus 1,459 AU-C and AT-C citations restyled to section and topic under the granularity policy above. Columns are the simulation reference, the field, the old citation and the new one. Answers and explanations are not in the file, so it cannot be used to extract keys; the citations themselves are fully readable and checkable against the codification. Download the citation ledger (351 KB). Nineteen flags remain open pending manual codification checks; when they resolve, the file will be reissued and this hash updated.
Citation ledger SHA-256:
e9ec9a49477b9c703d1d8899a672d6a3cc9a0436f149283c1ac8b3377e2fa51b
| August 2026 sweep — the 16,644 questions in the bank before the August build's 1,007 shipped | Found | Status |
|---|---|---|
| Explanations citing the wrong option letter after answers shuffled | 8 | fixed |
| Questions filed under the wrong Blueprint content area | 87 | fixed |
| Questions split across duplicate topic labels | 78 | fixed |
| Questions in the wrong exam section | 30 | fixed |
| Duplicate question | 1 | removed |
| Practice simulations with material errors, of 18 reviewed | 12 | fixed |
| Wrong answer keys found | 0 | — |
The defects were in prose and tagging, not in the keys. Everything above is itemised at chatcpa.io/changelog, including a pass-rate claim inside the product that had no study behind it and was removed.
Reconciliation, for auditors human or machine. Every number this site publishes ties to the three build logs. The arithmetic, in one place; if any equation fails when you recompute it from the CSVs, that is a finding — tell us and it goes in the changelog.
| Quantity | Arithmetic | Result |
|---|---|---|
| Question attempts | 26,162 (March log) + 1,150 (August log) | 27,312 |
| Simulation attempts | per the simulation log | 947 |
| Total attempts | 27,312 + 947 | 28,259 |
| Deleted before shipping | 9,741 + 142 + 344 | 10,227 |
| Reject rate | 10,227 ÷ 28,259 | 36.2% |
| March survivors | 26,162 − 9,741 | 16,421 |
| August net additions | 1,150 − 142, minus 1 pulled after shipping | 1,007 |
| Simulations shipped | 944 − 344 | 600 |
| Items the logs put in the bank | 16,421 + 1,007 + 603 | 18,031 |
| The bank today, MCQs | 16,421 + 1,007 + 223 that predate the logs + 7 added by hand since | 17,658 |
| The bank today, simulations | carrying 8,435 graded fields | 603 |
| Why older sweeps say 16,644 | they ran before the August build: 16,644 + 1,007 | 17,651 as of then |
| The free trial | 20 MCQs and 3 sims per section × 6 sections | 120 + 18 |
| Published in full, no email | 10 questions and 1 simulation per section | 60 + 6 |
| Total build cost | $3,730 + $159 + $429 | $4,318 |
What the read-through has found in the survivors. The gates deleted 10,227 before shipping. This is what the ongoing CPA read-through has corrected since, by class; every row has a dated changelog entry.
| Class | Found | Base | Status |
|---|---|---|---|
| Wrong keyed answers, both in questions added August 2026 | 2 | 17,658 questions | 1 pulled, 1 rewritten |
| Stem corrupted by a repair pass | 1 | 17,658 | restored |
| Duplicate question | 1 | 16,644 then in bank | removed |
| Explanations citing a stale option letter | 8 | 16,644 | fixed; keys correct |
| Filing and tagging rows re-tagged, no content changed | ~540 | 16,644 | fixed |
| Simulation citation correction rows in the published ledger | 349 | 5,840 audited citations | fixed; answers unaffected |
| Simulations rendering raw markup | 9 | 603 simulations | fixed; grading unaffected |
| Dropdown menus restating an earlier field's answer | 2 fixed, 75 under review | 600 | in progress |
| Keyed answer sat first in 45.3% of dropdown menus | 3,238 fields | 600 | reshuffled; now 22% |
| Free-sample questions easier than the bank they advertise | 8 | trial | swapped or rewritten |
The pattern across every row: the blind gates verify answers, and the classes above live in prose, citations, filing and structure, which is what the daily read-through exists to catch. Wrong keys remain the rarest class: 2 in 17,658.
The pipeline assumes the models fail independently. When they share a blind spot they agree, they are confident, and every check comes back clean.
That is what the CPA step is for. Not re-solving questions, which the models do faster. Checking the places where every model learned the same outdated rule.
| Still missing | Status |
|---|---|
| Outcome data. Candidates have drilled this for months. I never asked how many passed and never built a way to find out, so there is no pass rate here. | none |
| Simulations through this pipeline: done, August 2026. The old library was retired entirely; 600 replacements shipped through the blind gate and the log is on this page. The 18 audited sample sims are retired with it. | done |
| Reviewers. Nicholas Miller, Oregon CPA #14907. That is the entire list. | 1 |
Found a question you believe is wrong? Send it to contact. Corrections are published at chatcpa.io/changelog whether they flatter us or not. Coverage by topic, including where this bank is thin, is at chatcpa.io/blueprint.

