Test design

The evaluation used a fixed, stratified set of U.S. federal tax questions. Each question was answered twice by the same model: once without TaxMCP and once with TaxMCP connected. The paired design was intended to isolate the effect of access to retrieved primary authority.

Answers were graded against primary tax authority. The published homepage examples cover the post-OBBBA interest-limitation basis under IRC §163(j), the 2026 information-reporting threshold under §6041, and the bonus-depreciation rule under §168(k).

Metric definitions

Reported results

The TaxMCP-assisted run received a 9.8/10 mean overall grade versus 7.4 without retrieval, and was right or tied on 93% of questions. It scored 10/10 on the recent-law subset versus 4.8 without TaxMCP. The assisted answers were free of critical errors 96.7% of the time and stated evaluated figures and dates correctly 99.3% of the time.

How to read the results

This is a TaxMCP-run paired benchmark of one model run, not a third-party certification or a universal accuracy guarantee. It demonstrates the measured effect of TaxMCP retrieval on this question set. The complete prompt set, raw responses, grader materials, and model/version identity are not currently published, so independent reproduction will require the next release of benchmark materials.

Model behavior, source coverage, prompting, and tax law can change. That is why TaxMCP returns linked authority: every production answer can be checked against the underlying source by the professional responsible for the conclusion.

Building on the benchmark

The next edition is planned to name the tested model and version, publish the complete de-identified question set and grading rubric, repeat each condition to report variance, preserve raw outputs, and include an external tax practitioner in the review process. Those additions will make a compelling first result easier to reproduce and extend.