In October 2025, Deloitte’s Australian firm delivered a AU$440,000 report to the federal government, a review of the IT system the welfare department used to automate penalties. The report also contained a quote from a Federal Court judgment that the court never wrote, and footnotes to academic papers that did not exist. A university researcher caught it. Deloitte agreed to refund part of the fee, and a Big Four name spent a week as a cautionary headline.
That was a professional-services firm with a review process, not a solo preparer working late before a deadline. The fabrications shipped anyway.
For tax practitioners, the moment this stopped being someone else’s problem has a date: June 24, 2026.
The office that disciplines you just described your AI workflow
On that day the IRS Office of Professional Responsibility — the office that investigates Circular 230 violations and can suspend or disbar you from practice before the IRS — issued Alert 2026-19, “Introductory Guidelines for Responsible AI Use in Federal Tax Practice.”
Read who wrote it before you read what it says. This is not a vendor white paper or a CPE handout. It is the disciplinary body telling you, in advance, how it will read the existing rules when AI is in the loop.
The alert is blunt about the technology. It warns of “fabricated outputs (or, as commonly termed, ‘hallucinations’), bias, and lack of transparency.” Then it puts the burden in a single sentence: “it is incumbent on any tax professional using GAI to carefully review all documents crafted by the technology.”
The tool can draft; you remain accountable for what it produces. The rest of the alert maps that responsibility to existing rules.
Circular 230 doesn’t care that a machine made the mistake
The alert invents no new obligations. It maps AI onto the duties you already carry. Every one of these predates ChatGPT; the only new variable is the cause of the error.
| Circular 230 | Duty | What it means when AI drafts the work |
|---|---|---|
| § 10.22 | Diligence as to accuracy | Verify every fact, citation, and calculation the model produced before it reaches a client or the IRS. |
| § 10.35 | Competence | Understand the tool’s mechanics, limits, and failure modes — not only the law. |
| § 10.37 | Requirements for written advice | Base advice on assumptions and citations you have checked, not on the model’s confidence. |
| § 10.36 | Procedures to ensure compliance | Partners must set firm policy: staff training, approved tools, data-handling protocols, documentation. |
| § 10.27 | Fees | Do not bill for manual labor or time that was not actually spent, or double bill for AI-assisted work; the alert says whether a fee violates the rule depends on the facts. |
| § 10.51(a)(15) | Confidentiality | Use appropriately secured, firm-approved systems and follow the rules governing disclosure or use of tax return information. |
The confidentiality duty has teeth beyond Circular 230. Uploading return information to an unsecured tool can trigger IRC §7216, a criminal provision, and a civil penalty under §6713. Those penalties were written for the era of the fax machine. They read exactly the same way for a chatbot.
So is it malpractice?
Short version: the OPR alert does not say that using AI is itself malpractice. It says the practitioner’s existing duties of diligence, competence, confidentiality, and care in written advice continue to apply. Whether conduct amounts to malpractice depends on the facts and applicable law.
The line is the signature. The moment work goes in front of a client or gets filed with the IRS, §10.22 says you vouched for its accuracy and §10.37 says the advice rests on assumptions you actually checked. The model is not a practitioner. It cannot exercise diligence, and OPR will not accept it as the party who did. As the alert puts it, the “technology serves as a powerful tool, not a substitute for professional judgment,” and “final decisions must always rest with qualified professionals.”
Which sounds manageable — until you try to actually verify a general-purpose model’s answer.
The verification trap
The instruction is simple: check every citation. The problem is what you are checking it against.
When you ask ChatGPT for the Code section that governs a deduction, it does not look the section up. It predicts the most plausible next token, and a section number is a short string in a shape it has seen a million times. The number it hands back was generated, not retrieved. The same mechanism invents Tax Court cases that read perfectly and were never decided.
So “verify the output” quietly becomes a trap. There is no authoritative source behind the answer to check it against — only the model’s training, which is the very thing that produced the error. You end up checking a guess against your own memory, which is the work AI was supposed to save you.
The duty is to verify against primary authority. A tool that generates its own citations gives you nothing to verify against. That is the gap the alert names and does not close.
Close the gap: give the model a source it can’t fake
This is fixable, and the fix is not “use AI less.” It is to change where the answer comes from. If the citation is retrieved from an index of real authority and handed to the model — rather than generated by it — a fabricated section or a phantom case never appears, because it was never in the source to copy.
That is what taxmcp.io is built to do. Retrieved results carry a Citation URL to the primary source — the IRC section, Treasury regulation, IRS publication, ruling, or (on Pro+) the Tax Court opinion — that you can open directly. The work demanded by §10.22 becomes easier when the claimed authority is one click away instead of something you must reconstruct from the model’s memory.
TaxMCP does not require a client return or supporting documents, and it does not persist MCP tool arguments in its application database. Search text is processed through OpenAI’s API for embeddings and, when enabled, reranking under OpenAI’s API data controls. That can help a firm minimize the information sent to an additional research service, but it does not make the surrounding Claude or ChatGPT workflow compliant by itself. Information a practitioner enters into the AI client remains governed by that provider’s terms, the firm’s policies, and the practitioner’s applicable obligations. See the full security and data-handling description.
It does not take you out of the loop. OPR is explicit that nothing can. It gives your judgment something real to land on.
The argument worth having now
The alert is clear about responsibility: the practitioner remains responsible when AI gets the law wrong. The practical question is how to make primary-source review routine enough to perform on every engagement.
See how taxmcp grounds every answer in primary authority →
Or start with the broader question the alert forces: can you actually trust AI for tax research?