Methodology
How the questions get made, and what gets deleted.
Anyone can ask a model for a hundred questions. Finding its errors is the product. Here is the process, the numbers, and the part it cannot catch.
| # | Step | Who | Sees the answer | May edit | Produces |
|---|---|---|---|---|---|
| 1 | Write | Model A | writes it | yes | One question, four options, four explanations |
| 2 | Judge | Model B | yes | no | Accept, or reject with a reason |
| 3 | Blind solve | Model C | no | no | One letter |
| 4 | Blind solve | Model D | no | no | One letter |
| 5 | Compare | Arithmetic | — | no | Ship, or delete |
| 6 | Sign off, then keep reading | CPA | yes | yes | Correlated-error check, ongoing |
On step 6, because the honest version matters. The timestamps in the log cover generation only, a 16-day window in March. Review is not a stage that finished when the build did. I read no fewer than 50 questions a day, every day, and I am still doing it five months later. That is several thousand questions since the build and counting, which is why the changelog keeps getting entries rather than stopping at launch. What I have not done, and will not claim, is personally re-solve all 16,421 at build time. Nobody could.
Each model is from a different family. Blind solvers receive the stem and four options only: no key, no explanation, no hint, no topic label. A rejected question is never repaired, it is deleted and written again from scratch, because editing a flawed question preserves the shape that made it flawed. Every question is commissioned against one specific Blueprint point, not a topic, so questions from the same topic cannot collapse into each other.
| Condition | Required | Fails to |
|---|---|---|
| Blind solver C vs blind solver D | same letter | delete |
| Blind letter vs recorded answer | match | review by hand |
| Answerable as written | yes, both | delete |
| Blocking defect reported | none, both | delete |
| Second defensible answer | none, both | delete |
| Judge verdict | accept | regenerate once |
| Model confidence score | not used | — |
Confidence is deliberately excluded. In this build a model returned 99 out of 100 on a question with two defensible answers. A confidence floor would also delete good hard questions, which score lower precisely because they are discriminating.
| Why 9,741 questions were deleted | Gate 1 | Gate 2 |
|---|---|---|
| Missing a fact needed to decide the answer | 2,761 | 614 |
| Technical error | 2,241 | 1,683 |
| Two defensible answers | 1,082 | 440 |
| Explanation conflicts with the key or the stem | 634 | 445 |
| Correct only under superseded guidance | 486 | 191 |
| Stem contradicts itself | 404 | 719 |
| Misleading wording | 302 | 368 |
| Unsafe date or threshold | 127 | — |
| Other trust risk | 1 | 127 |
| Rejected by this gate | 8,038 | 4,587 |
The two gates fail differently. Gate one's most common finding was a missing decisive fact; gate two's was a technical error. They overlapped on only 2,884 questions.
| The build, 15 to 30 March 2026 | Count | Share |
|---|---|---|
| Questions generated | 26,162 | — |
| Questions deleted | 9,741 | 37.2% |
| Questions kept | 16,421 | 62.8% |
| Where the money went | Total | Per attempt | Share |
|---|---|---|---|
| Writing the questions | $1,425 | $0.054 | 38% |
| Checking them: judge, two independent gates, orchestration | $2,305 | $0.089 | 62% |
| Total | $3,730 | $0.143 | 100% |
| Cost per question that survived | — | $0.227 | — |
| Earlier build, written off entirely | $2,000 | — | — |
Generating the questions was 38 percent of the cost. Checking them was the other 62. That ratio is the entire difference between this and a bank someone produced by asking a model for a hundred questions two hundred times. Five and a half cents makes a question. Twenty-three cents puts one in this bank, and the gap is the 9,741 that were deleted.
The full log is published. Every one of the 26,162 attempts, with its timestamp, section, topic, both gate verdicts, the reject code, and what it cost. Download the CSV (4.8 MB), or read the summary JSON if you just want the totals.
CSV SHA-256:
dca855fd123ac340490d666b834fd5ed800a8f9fb56b02464a8cc9f2e6df553a
Published so you can confirm the file you downloaded is the file we published, and so we cannot
quietly swap it later.
This log covers the March 2026 build only, from 15 to 30 March. References are assigned when a question is loaded into the bank, not when it is generated, so the log records what was written and judged rather than what was ultimately loaded. The bank holds 16,644 questions; 16,421 of them came from this build and the rest predate it.
| August 2026 sweep, all 16,644 questions | Found | Status |
|---|---|---|
| Explanations citing the wrong option letter after answers shuffled | 8 | fixed |
| Questions filed under the wrong Blueprint content area | 87 | fixed |
| Questions split across duplicate topic labels | 78 | fixed |
| Questions in the wrong exam section | 30 | fixed |
| Duplicate question | 1 | removed |
| Practice simulations with material errors, of 18 reviewed | 12 | fixed |
| Wrong answer keys found | 0 | — |
The defects were in prose and tagging, not in the keys. Everything above is itemised at chatcpa.io/changelog, including a pass-rate claim inside the product that had no study behind it and was removed.
The pipeline assumes the models fail independently. When they share a blind spot they agree, they are confident, and every check comes back clean.
That is what the CPA step is for. Not re-solving questions, which the models do faster. Checking the places where every model learned the same outdated rule.
| Still missing | Status |
|---|---|
| Outcome data. Candidates have drilled this for months. I never asked how many passed and never built a way to find out, so there is no pass rate and no testimonials here. | none |
| Simulations through this same pipeline. The 18 in the free sample were audited and fixed. The full library has not been. | not done |
| Reviewers. Nicholas Miller, Oregon CPA #14907. That is the entire list. | 1 |
Found a question you believe is wrong? Send it to contact. Corrections are published at chatcpa.io/changelog whether they flatter us or not. Coverage by topic, including where this bank is thin, is at chatcpa.io/blueprint.

