Plenty of organizations have AI policies promising human review and responsible use. Attorney and CPA Justin Kavalir argues those statements are only assertions and a recent federal case shows what happens when one is tested and no evidence of the promised oversight can be produced.
An increasing number of organizations have policies on AI. Many contain some version of a statement calling for responsible AI use and claiming AI systems are subject to human oversight. Often, this includes language that human review or verification of AI output is required.
These policy statements are assertions, but they are only the beginning of governance. They raise basic questions. What outputs are subject to human review? Which humans perform the review and by what standard? What record is created when this review or verification occurs? When does the review or verification occur?
A recent Rule 11 federal sanctions order from the Western District of Tennessee is instructive here, highlighting the type of evidence of responsible AI use that organizations may be asked to produce. The case is not significant because of AI hallucinations, but it is worthy of analysis because of what evidence the court expected to be produced.
The Reaves case
In Reaves Law Firm v. Baker Donelson, et al., the defendants alleged that a motion filed by Reaves Law Firm contained the hallmarks of generative AI use by including arguments unsupported by the cited cases and direct quotes that did not exist. Reaves responded with two additional filings. The defendants asserted the two additional filings contained more generative AI hallucinations. The court, finding sufficient support for the defendants’ allegations, issued a show-cause order requiring Reaves to confirm whether the cases exist; whether the cases support the propositions for which they were offered; and whether quotes exist in the cases cited. Going beyond the determination of whether the cases, propositions or quotes exist or are supported as cited, the court ordered Reaves to describe its process, mandating the firm “identify what steps, if any, it had taken prior to citing to the cases to verify their existence.”
Reaves failed to provide the requested evidence of verification. The firm’s response included an internal email sent from one person to two others. The subject line was “Mandatory Ethical AI Training & Reporting Protocols for All Staff.” However, Reaves did not demonstrate the email was shared across the firm or that any training actually took place. The firm also indicated its general counsel was no longer employed there and that “internal filing supervision has been restructured.” The court observed that the firm’s namesake signed the pleadings himself and noted that laying blame on departed personnel did not relieve Reaves Law Firm of responsibility, sanctioning the firm under Rule 11.
Read closely, the inability of Reaves to produce any meaningful evidence to the court demonstrating verification became part of the failure by exposing the apparent absence of a process. Had the firm been able to demonstrate contemporaneous records of its verification processes, it could have answered the court’s central inquiry. In this case, no such records were produced. To be clear, records would not have made the citations real, but a firm that can show a documented verification process where personnel failed stands categorically different from a firm that can show neither. The firm could have provided an affidavit as evidence of its verification processes, but the affidavit would have been based on memory. Changing personnel and restructured processes made that reconstruction more difficult and less reliable.
What is also worth noting is that this order does not impose a new obligation unique to AI. Reaves was sanctioned under Rule 11, which has been on the books for nearly 90 years. As this case demonstrates, AI did not change an existing duty, but it highlights how easily a task that carries a duty can be outsourced to an AI system and how hard it is to demonstrate appropriate oversight after the fact.
A framework for the gap
The order maps to four capabilities an organization should possess for effective AI governance. An organization should define what decisions require human review, record the results of this review, assign ownership of the review to a named person and perform oversight of the reviews to guard against missed signals. This can be distilled to four practical actions: Define, record, own and guard.
- Define requires an organization to determine where AI systems influence consequential decisions that require human review. From the opinion, Reaves never meaningfully answered the question of AI’s role in its court pleadings at issue. The filing of pleadings in court is a consequential decision. It carries consequences to the attorney who signs the pleading, to the attorney’s firm and to the client.
- Record requires an organization to create evidence of what the AI system produced and how the human processed that output, including what was reviewed, what judgment was exercised and what conclusion was reached. The opinion does not indicate the production of any records to demonstrate the verification steps the court sought in this case. The firm offered a general policy or protocol but not evidence that it was followed to verify cases or citations.
- Own requires a named human to be assigned to a decision for accountability. In this case, the firm attempted to impose accountability on a departed general counsel; however, the court attributed accountability to the firm jointly with the attorney whose name appeared on the pleadings.
- Guard requires organizations to monitor their governance over time and across decisions to determine whether escalation, modification or suspension of a process is required. Here, the firm received a signal that should have resulted in a review of its processes following the first allegations of unsupported or hallucinated citations. Instead, the firm proceeded to file two additional pleadings with the same defects as the first pleading. A warning signal arrived, and the firm did not stop, review or correct its processes.
Taken together, a firm that can demonstrate these four functions under pressure with a contemporaneous record is in a fundamentally different position than Reaves when the show cause order arrived.
The practical lesson
The practical lesson of this case is not that organizations should stop using AI or abandon AI policies. AI policy statements requiring review of AI output are not worthless, but they may be incomplete. These statements may state an expectation without having the underlying infrastructure in place to evidence compliance. A general counsel, risk officer or a compliance leader evaluating an AI governance program should go beyond the policy document and ask whether it can produce evidence of compliance on demand.
This distinction between a governance program consisting of a policy document vs. a governance program that can produce evidence of compliance will not be confined to legal filings. Any organization making a claim that it maintains human-in-the-loop or practices responsible AI use is making an assertion, or, potentially, a representation. This assertion can be tested by opposing counsel, regulators, insurers, auditors, boards or by shareholders. The question worth asking now, before a show-cause order or its equivalent arrives, is simple: If asked tomorrow to prove it, what would there be to produce?


Justin Kavalir is an attorney, CPA and a university general counsel. He is the founder of the Judgment Assurance Institute. 







