Methodology

How the questions get made, and what gets deleted.

Anyone can ask a model for a hundred questions. Finding its errors is the product. Here is the process, the numbers, and the part it cannot catch.

#StepWhoSees the answerMay editProduces
1WriteModel Awrites ityesOne question, four options, four explanations
2JudgeModel ByesnoAccept, or reject with a reason
3Blind solveModel CnonoOne letter
4Blind solveModel DnonoOne letter
5CompareArithmeticnoShip, or delete
6Sign off, then keep readingCPAyesyesCorrelated-error check, ongoing

On step 6, because the honest version matters. The timestamps in the log cover generation only, a 16-day window in March. Review is not a stage that finished when the build did. I read no fewer than 50 questions a day, every day, and I am still doing it five months later. That is several thousand questions since the build and counting, which is why the changelog keeps getting entries rather than stopping at launch. What I have not done, and will not claim, is personally re-solve all 16,421 at build time. Nobody could.

Each model is from a different family. Blind solvers receive the stem and four options only: no key, no explanation, no hint, no topic label. A rejected question is never repaired, it is deleted and written again from scratch, because editing a flawed question preserves the shape that made it flawed. Every question is commissioned against one specific Blueprint point, not a topic, so questions from the same topic cannot collapse into each other.

ConditionRequiredFails to
Blind solver C vs blind solver Dsame letterdelete
Blind letter vs recorded answermatchreview by hand
Answerable as writtenyes, bothdelete
Blocking defect reportednone, bothdelete
Second defensible answernone, bothdelete
Judge verdictacceptregenerate once
Model confidence scorenot used

Confidence is deliberately excluded. In this build a model returned 99 out of 100 on a question with two defensible answers. A confidence floor would also delete good hard questions, which score lower precisely because they are discriminating.

Why 9,741 questions were deletedGate 1Gate 2
Missing a fact needed to decide the answer2,761614
Technical error2,2411,683
Two defensible answers1,082440
Explanation conflicts with the key or the stem634445
Correct only under superseded guidance486191
Stem contradicts itself404719
Misleading wording302368
Unsafe date or threshold127
Other trust risk1127
Rejected by this gate8,0384,587

The two gates fail differently. Gate one's most common finding was a missing decisive fact; gate two's was a technical error. They overlapped on only 2,884 questions.

1,703 questions were caught by the second gate alone. Gate one had already approved them. Run one reviewer instead of two and all 1,703 would be in the bank right now. That is the measured value of checking the same question twice, independently, and it is the reason both gates exist rather than one careful one.
The build, 15 to 30 March 2026CountShare
Questions generated26,162
Questions deleted9,74137.2%
Questions kept16,42162.8%
Where the money wentTotalPer attemptShare
Writing the questions$1,425$0.05438%
Checking them: judge, two independent gates, orchestration$2,305$0.08962%
Total$3,730$0.143100%
Cost per question that survived$0.227
Earlier build, written off entirely$2,000

Generating the questions was 38 percent of the cost. Checking them was the other 62. That ratio is the entire difference between this and a bank someone produced by asking a model for a hundred questions two hundred times. Five and a half cents makes a question. Twenty-three cents puts one in this bank, and the gap is the 9,741 that were deleted.

The full log is published. Every one of the 26,162 attempts, with its timestamp, section, topic, both gate verdicts, the reject code, and what it cost. Download the CSV (4.8 MB), or read the summary JSON if you just want the totals.

CSV SHA-256: dca855fd123ac340490d666b834fd5ed800a8f9fb56b02464a8cc9f2e6df553a
Published so you can confirm the file you downloaded is the file we published, and so we cannot quietly swap it later.

This log covers the March 2026 build only, from 15 to 30 March. References are assigned when a question is loaded into the bank, not when it is generated, so the log records what was written and judged rather than what was ultimately loaded. The bank holds 16,644 questions; 16,421 of them came from this build and the rest predate it.

August 2026 sweep, all 16,644 questionsFoundStatus
Explanations citing the wrong option letter after answers shuffled8fixed
Questions filed under the wrong Blueprint content area87fixed
Questions split across duplicate topic labels78fixed
Questions in the wrong exam section30fixed
Duplicate question1removed
Practice simulations with material errors, of 18 reviewed12fixed
Wrong answer keys found0

The defects were in prose and tagging, not in the keys. Everything above is itemised at chatcpa.io/changelog, including a pass-rate claim inside the product that had no study behind it and was removed.

The pipeline assumes the models fail independently. When they share a blind spot they agree, they are confident, and every check comes back clean.

A real example from this build. A question asked by how much a Section 179 deduction is reduced for a taxpayer placing $3,250,000 of property in service. Both blind solvers agreed with the key. Both were confident. The judge passed it. All three were applying a phase-out threshold that legislation had raised, so under current law that taxpayer has no reduction at all and the correct answer was not among the four options. Caught in human review, deleted.

That is what the CPA step is for. Not re-solving questions, which the models do faster. Checking the places where every model learned the same outdated rule.
Still missingStatus
Outcome data. Candidates have drilled this for months. I never asked how many passed and never built a way to find out, so there is no pass rate and no testimonials here.none
Simulations through this same pipeline. The 18 in the free sample were audited and fixed. The full library has not been.not done
Reviewers. Nicholas Miller, Oregon CPA #14907. That is the entire list.1

Found a question you believe is wrong? Send it to contact. Corrections are published at chatcpa.io/changelog whether they flatter us or not. Coverage by topic, including where this bank is thin, is at chatcpa.io/blueprint.