CPAPreference Bench
CPAPreferenceBench: practitioner adjudication on 1,860 complex personal tax-review findings.
Introduction
When an AI independently second-reviews a professionally prepared Form 1040 engagement which of its Findings across accuracy, completeness, and tax efficiency would a CPA actually want to use?
Eight frontier models reviewed the same 25 real-world tax engagements. A team of CPAs then blindly adjudicated 1,860 Findings for factual support, repair needed, and practitioner value.
CPAPreference Bench leaderboard
Claude Opus 5[max]5.00
95% interval 4.00–6.12Claude Fable 5[max]4.08
95% interval 3.28–4.92Kimi K3[max]4.08
95% interval 3.40–4.76Muse Spark 1.1[xhigh]3.88
95% interval 3.16–4.68Grok 4.5[high]3.76
95% interval 3.04–4.48GPT-5.6 Sol[max]3.64
95% interval 3.00–4.36Gemini 3.6 Flash[high]3.32
95% interval 2.72–3.92Inkling[max]3.00
95% interval 2.36–3.64
Overview
Tax review is long-context work. The useful signal may sit across a filed return, prior-year return, brokerage statements, retirement documents, handwritten notes, and other messy supporting material. A plausible sentence is not enough: a Finding must be grounded in the Binder and useful to the practitioner responsible for the return.
Each model ran three independent reviews—Accuracy for return-to-document discrepancies, Completeness for unexplained year-over-year changes, and Tax Efficiency for document-grounded planning opportunities. The models could inspect original page images and OCR, search a frozen set of IRS materials, execute Python, and maintain a durable notepad.
The sample taxpayers are over the applicable age threshold, report IRA distributions, and have a recurring pattern of charitable gifts. The sample records do not indicate that any gifts were paid directly from an IRA.
Interactive example of CPA adjudication for Tax Efficiency section.
Methodology
Every Engagement produced one Run per model and one Category Execution per review category. Models received the same approved Binder, the same current- and prior-year Law Packs, the same tools, and category-specific instructions. Model identity was hidden behind randomized Review A–H positions while CPAs adjudicated Findings.
The headline measure is Valuable Findings per Engagement. A Finding is Valuable only when the CPA reviewers marked it Supported and assigned Practitioner Value 2 or 3. Failed category executions remain part of the denominator and contribute zero Findings.
Three review categories
Reconciles the filed return against current-year support and surfaces material discrepancies.
Compares current and prior returns for material continuity questions that need follow-up.
Identifies concrete federal planning opportunities supported by the client record.
Agent environment
Models could list, search, read, and view Binder documents; separately inspect a frozen IRS Law Pack; run network-isolated Python; maintain a persistent notepad; and submit only their assigned category schema. Context was compacted only after a model crossed 80% of its configured window.
CPA adjudication rubric
Is the core claim established by the available evidence?
Can it be used as written, or does it need minor, major, or complete replacement?
Does it add no value, modest value, material value, or unusually high value?
Model Configuration
All eight models used the highest supported reasoning control selected for their provider. Those labels are provider-native settings, not comparable reasoning-token budgets. No model received an application temperature override.
| Model | Provider · reasoning | Context window |
|---|---|---|
| Claude Fable 5 | Anthropic · max | 1,000,000 tokens |
| Claude Opus 5 | Anthropic · max | 1,000,000 tokens |
| Gemini 3.6 Flash | Google · high | 1,048,576 tokens |
| GPT-5.6 Sol | OpenAI · max | 1,050,000 tokens |
| Grok 4.5 | xAI · high | 500,000 tokens |
| Inkling | Thinking Machines · max | 1,040,000 tokens |
| Kimi K3 | Kimi · max | 1,048,576 tokens |
| Muse Spark 1.1 | Meta · xhigh | 1,048,576 tokens |
Results
Models differed both in how many useful Findings they surfaced and in how densely that value appeared in their submissions. The table keeps yield, volume, and support separate rather than collapsing them into a composite score; execution reliability is reported as a cohort statistic.
| 12.88 | 5.004.00–6.12 95% CI | 38.8% | 96.0% | 1 | 99.7% | |
| 8.72 | 4.083.28–4.92 95% CI | 46.8% | 96.0% | 1 | 98.6% | |
| 8.48 | 4.083.40–4.76 95% CI | 48.1% | 100.0% | 1 | 99.1% | |
| 9.32 | 3.883.16–4.68 95% CI | 41.6% | 100.0% | 1 | 100.0% | |
| 7.44 | 3.763.04–4.48 95% CI | 50.5% | 96.0% | 1 | 99.5% | |
| 8.68 | 3.643.00–4.36 95% CI | 41.9% | 96.0% | 0 | 98.6% | |
| 10.60 | 3.322.72–3.92 95% CI | 31.3% | 96.0% | 1 | 99.6% | |
| 8.28 | 3.002.36–3.64 95% CI | 36.2% | 88.0% | 0 | 98.1% |
25 blind Engagement reviews per model. 1,860 adjudicated Findings. 598/600 Category Executions succeeded. Valuable Yield is the share of submitted Findings that were both Supported and rated Practitioner Value 2 or 3. Exceptional Findings received Practitioner Value 3. Support excludes Findings marked as duplicates.
Limitations
These results cover the currently completed subset: 25 tax-year 2025 Engagements from one partner firm. The benchmark does not contain an exhaustive expert-authored inventory of every possible Finding, so it measures confirmed yield and quality—not absolute recall.
Two of 600 Category Executions failed; under the system-yield definition, they remain in the denominator and contribute zero Findings.
Acknowledgements
We thank the partner accounting firm that supplied professionally prepared, real-world tax documents and the team of CPAs that reviewed and adjudicated the model Findings.
Get in touch
Run your agent against the benchmark, get expert data, chat with the team