Test design
The evaluation used a fixed, stratified set of U.S. federal tax questions. Each question was answered twice by the same model: once without TaxMCP and once with TaxMCP connected. The paired design was intended to isolate the effect of access to retrieved primary authority.
Answers were graded against primary tax authority. The published homepage examples cover the post-OBBBA interest-limitation basis under IRC §163(j), the 2026 information-reporting threshold under §6041, and the bonus-depreciation rule under §168(k).
Metric definitions
- Overall accuracy grade: the mean question-level score on a 0–10 scale.
- Free of critical errors: the share of answers without an error serious enough to change or materially undermine the conclusion.
- Exact figures and dates correct: the share of evaluated numeric thresholds, rates, and dates stated correctly.
- Right or tied: the share of paired questions where the TaxMCP-assisted answer scored at least as well as the unassisted answer.
- Recent-law score: the mean 0–10 grade on the recent-law subset.
Reported results
The TaxMCP-assisted run received a 9.8/10 mean overall grade versus 7.4 without retrieval, and was right or tied on 93% of questions. It scored 10/10 on the recent-law subset versus 4.8 without TaxMCP. The assisted answers were free of critical errors 96.7% of the time and stated evaluated figures and dates correctly 99.3% of the time.
How to read the results
This is a TaxMCP-run paired benchmark of one model run, not a third-party certification or a universal accuracy guarantee. It demonstrates the measured effect of TaxMCP retrieval on this question set. The complete prompt set, raw responses, grader materials, and model/version identity are not currently published, so independent reproduction will require the next release of benchmark materials.
Model behavior, source coverage, prompting, and tax law can change. That is why TaxMCP returns linked authority: every production answer can be checked against the underlying source by the professional responsible for the conclusion.
Building on the benchmark
The next edition is planned to name the tested model and version, publish the complete de-identified question set and grading rubric, repeat each condition to report variance, preserve raw outputs, and include an external tax practitioner in the review process. Those additions will make a compelling first result easier to reproduce and extend.