CPAPreference Bench

CPAPreferenceBench: practitioner adjudication on 1,860 complex personal tax-review findings.

Introduction

When an AI independently second-reviews a professionally prepared Form 1040 engagement which of its Findings across accuracy, completeness, and tax efficiency would a CPA actually want to use?

Eight frontier models reviewed the same 25 real-world tax engagements. A team of CPAs then blindly adjudicated 1,860 Findings for factual support, repair needed, and practitioner value.

CPAPreference Bench leaderboard

  1. Claude Opus 5[max]5.00
    95% interval 4.006.12
  2. Claude Fable 5[max]4.08
    95% interval 3.284.92
  3. Kimi K3[max]4.08
    95% interval 3.404.76
  4. Muse Spark 1.1[xhigh]3.88
    95% interval 3.164.68
  5. Grok 4.5[high]3.76
    95% interval 3.044.48
  6. GPT-5.6 Sol[max]3.64
    95% interval 3.004.36
  7. Gemini 3.6 Flash[high]3.32
    95% interval 2.723.92
  8. Inkling[max]3.00
    95% interval 2.363.64
Mean CPA-confirmed Valuable Findings per Engagement. Whiskers show 95% Engagement-bootstrap intervals; this is one outcome measure, not a composite score.

Overview

Tax review is long-context work. The useful signal may sit across a filed return, prior-year return, brokerage statements, retirement documents, handwritten notes, and other messy supporting material. A plausible sentence is not enough: a Finding must be grounded in the Binder and useful to the practitioner responsible for the return.

Each model ran three independent reviews—Accuracy for return-to-document discrepancies, Completeness for unexplained year-over-year changes, and Tax Efficiency for document-grounded planning opportunities. The models could inspect original page images and OCR, search a frozen set of IRS materials, execute Python, and maintain a durable notepad.

Qualified charitable distribution planningMedium
Easy

The sample taxpayers are over the applicable age threshold, report IRA distributions, and have a recurring pattern of charitable gifts. The sample records do not indicate that any gifts were paid directly from an IRA.

Support

Interactive example of CPA adjudication for Tax Efficiency section.

Methodology

Every Engagement produced one Run per model and one Category Execution per review category. Models received the same approved Binder, the same current- and prior-year Law Packs, the same tools, and category-specific instructions. Model identity was hidden behind randomized Review A–H positions while CPAs adjudicated Findings.

The headline measure is Valuable Findings per Engagement. A Finding is Valuable only when the CPA reviewers marked it Supported and assigned Practitioner Value 2 or 3. Failed category executions remain part of the denominator and contribute zero Findings.

Three review categories
Accuracy

Reconciles the filed return against current-year support and surfaces material discrepancies.

Completeness

Compares current and prior returns for material continuity questions that need follow-up.

Efficiency

Identifies concrete federal planning opportunities supported by the client record.

Agent environment

Models could list, search, read, and view Binder documents; separately inspect a frozen IRS Law Pack; run network-isolated Python; maintain a persistent notepad; and submit only their assigned category schema. Context was compacted only after a model crossed 80% of its configured window.

CPA adjudication rubric
Support

Is the core claim established by the available evidence?

Repair

Can it be used as written, or does it need minor, major, or complete replacement?

Practitioner Value

Does it add no value, modest value, material value, or unusually high value?

Model Configuration

All eight models used the highest supported reasoning control selected for their provider. Those labels are provider-native settings, not comparable reasoning-token budgets. No model received an application temperature override.

Model provider, reasoning setting, and context window
ModelProvider · reasoningContext window
Claude Fable 5Anthropic · max1,000,000 tokens
Claude Opus 5Anthropic · max1,000,000 tokens
Gemini 3.6 FlashGoogle · high1,048,576 tokens
GPT-5.6 SolOpenAI · max1,050,000 tokens
Grok 4.5xAI · high500,000 tokens
InklingThinking Machines · max1,040,000 tokens
Kimi K3Kimi · max1,048,576 tokens
Muse Spark 1.1Meta · xhigh1,048,576 tokens

Results

Models differed both in how many useful Findings they surfaced and in how densely that value appeared in their submissions. The table keeps yield, volume, and support separate rather than collapsing them into a composite score; execution reliability is reported as a cohort statistic.

CPAPreference Bench model results
Claude Opus 512.885.004.006.12 95% CI38.8%96.0%199.7%
Claude Fable 58.724.083.284.92 95% CI46.8%96.0%198.6%
Kimi K38.484.083.404.76 95% CI48.1%100.0%199.1%
Muse Spark 1.19.323.883.164.68 95% CI41.6%100.0%1100.0%
Grok 4.57.443.763.044.48 95% CI50.5%96.0%199.5%
GPT-5.6 Sol8.683.643.004.36 95% CI41.9%96.0%098.6%
Gemini 3.6 Flash10.603.322.723.92 95% CI31.3%96.0%199.6%
Inkling8.283.002.363.64 95% CI36.2%88.0%098.1%

25 blind Engagement reviews per model. 1,860 adjudicated Findings. 598/600 Category Executions succeeded. Valuable Yield is the share of submitted Findings that were both Supported and rated Practitioner Value 2 or 3. Exceptional Findings received Practitioner Value 3. Support excludes Findings marked as duplicates.

Limitations

These results cover the currently completed subset: 25 tax-year 2025 Engagements from one partner firm. The benchmark does not contain an exhaustive expert-authored inventory of every possible Finding, so it measures confirmed yield and quality—not absolute recall.

Two of 600 Category Executions failed; under the system-yield definition, they remain in the denominator and contribute zero Findings.

Acknowledgements

We thank the partner accounting firm that supplied professionally prepared, real-world tax documents and the team of CPAs that reviewed and adjudicated the model Findings.

Get in touch

Run your agent against the benchmark, get expert data, chat with the team