Judgment Assurance:
AI Governance Liability Tracker
Judgment Assurance is a decision governance framework for making human judgment in AI-mediated consequential decisions explicit, evidenced, and institutionally accountable. The AI Governance Liability Tracker is a research resource within the Judgment Assurance framework.The Liability Tracker monitors litigation, regulatory actions, enforcement activity, insurance developments, and other events that shape institutional accountability for AI-mediated consequential decisions. It provides a periodically updated evidence base that helps organizations understand how governance expectations are evolving in practice and why documented human judgment increasingly matters.Each entry is summarized for its relevance to Judgment Assurance: the distinction between claimed oversight and evidenced oversight.The AI Governance Liability Tracker supports the Judgment Assurance framework by documenting the legal, regulatory, and institutional evidence surrounding accountability for AI-mediated consequential decisions.
© 2026 Judgment Assurance Institute, LLCAll rights reserved.
Anderson v. Microsoft Corp. - Complaint
W.D. Wash., No. 2:26-cv-02281 · Filed June 30, 2026
What Happened
Plaintiff filed a shareholder derivative complaint alleging breach of fiduciary duty in connection with Microsoft's training of its AI systems on copyrighted materials without proper licenses, and with representations to shareholders that its AI development complied with copyright law and that the Board maintained active oversight of AI strategy.
Why It Matters
What makes this complaint notable is how heavily it relies on the board's own public representations about that oversight. Three consecutive years of Proxy Statements affirmatively stated that the Board "maintains direct oversight over AI strategy risk" and that its Environmental, Social, and Public Policy Committee actively received and responded to management updates on AI governance. Having claimed that role publicly, the board is now alleged to have held it during a period when the underlying misconduct, and a specific red flag (the June 2025 Bartz v. Anthropic ruling on training-data provenance), went unaddressed. The fiduciary breach theory follows closely from the board's own assertion: you told shareholders this was your responsibility, and it happened on your watch.Whether Microsoft's actual governance record supports or undermines that claim is a question for discovery and further pleadings. The company may have exactly the record needed to defeat the claim at the pleading stage or later. But the case is a clean illustration of the exposure created once an organization asserts oversight publicly: the assertion becomes the standard you're measured against, whether or not you're prepared to be measured.
JA Relevance
Judgment Assurance is not a data-governance or IP compliance tool, and nothing here suggests it would have prevented the underlying conduct alleged in this complaint. The decision to train on unlicensed data is a sourcing and legal-compliance question, not an operational reliance failure. What the case does illustrate, at a structural level, is the risk inherent in any public claim of active oversight: it becomes the standard a plaintiff will use to measure your conduct after the fact. That dynamic exists within workflows and individual decisions that rely on AI output, which is where JA is built to provide evidence of the asserted oversight.
SEIU Pension Plan Master Trust v. Narayen, et al. (Adobe) - Complaint
N.D. Cal., No. 3:26-cv-03521 · Filed April 24, 2026
What Happened
Plaintiff filed a stockholder derivative complaint against Adobe directors and senior officers alleging breach of fiduciary duty, corporate waste, and violations of Sections 10(b) and 14(a) of the Securities Exchange Act. The complaint alleges that Adobe's SlimLM small language models were trained on SlimPajama, a copied, cleaned, and deduplicated derivative of RedPajama, which in turn allegedly incorporated Books3, a dataset of roughly 196,000 pirated books, and that Adobe personnel selected and used this dataset despite known legal risk.The complaint centers on Adobe's public filings. Plaintiff alleges that Adobe's 2023 through 2025 annual reports and 2024 and 2025 proxy statements represented that Adobe's Firefly models were commercially safe and trained only on licensed or public-domain content, and that Adobe's 2026 proxy statement, filed after copyright class actions were already pending, quietly omitted that licensing language. Plaintiff characterizes the omission as a tacit admission that the earlier statements were false when made.
Why It Matters
This is not a simple AI copyright case. Its significance is that plaintiff attempts to convert Adobe's own AI assurance language into the basis for fiduciary-duty, securities, disclosure, and corporate-waste claims.The complaint's theory depends heavily on the specificity of Adobe's prior public statements. Adobe allegedly did not merely say that it cared about responsible AI or creator rights. Plaintiff alleges that Adobe made concrete, checkable claims about how certain AI models were trained, including claims that training data came from licensed, openly licensed, or public-domain sources and that resulting outputs were commercially safe.That distinction matters. General AI ethics language may be difficult to test. A factual representation about training-data sources is different. Once a company states that an AI system was built using particular categories of authorized content, that statement becomes measurable against the actual development record. If later litigation alleges that the system was trained on unauthorized or pirated material, the company's own assurance language can become the benchmark against which its conduct is judged.The 2026 proxy allegation sharpens that point further. Plaintiff highlights not just Adobe's prior AI assurance language, but Adobe's later removal of that language after copyright litigation had already been filed. Whether that omission actually supports an inference of falsity is a question for further pleading and, eventually, discovery. But the pleading shows how a company's own retreat from specific AI language, not just the language itself, can become part of the evidentiary record against it.
JA Relevance
This case is adjacent to Judgment Assurance rather than a direct human-oversight case. The complaint does not allege that a human reviewer rubber-stamped an AI-generated recommendation, failed to exercise override authority, or lacked an adequate decision record in an operational workflow. It concerns training-data provenance, public AI assurance statements, securities disclosures, and board-level oversight.Judgment Assurance is not a data-governance, model-training, or IP-compliance tool. It would not determine whether a dataset was properly licensed, whether a model was trained on copyrighted material, or whether Adobe's public statements were accurate. Those are legal, sourcing, compliance, and disclosure-control questions outside JA's core function.The case is relevant because it illustrates a broader evidentiary principle that is central to JA: specific AI assurances become litigation artifacts. When an organization makes a factual claim about how an AI system is built, governed, reviewed, or controlled, that claim may later become the standard against which the organization is measured.The same dynamic applies to human oversight claims. Many organizations already say that AI only assists, that humans remain in the loop, that trained personnel review AI outputs, or that final decisions are made by accountable human decision-makers. Those statements are not risk-free. If challenged, the organization may need to show what the AI produced, what the human saw, what the human did with it, why the human accepted, rejected, modified, or escalated the output, and who owned the final decision.JA addresses that narrower problem. It creates contemporaneous evidence of institutional judgment in AI-mediated decision workflows, rather than leaving the organization to defend a generalized assurance after the fact. The lesson is not that Judgment Assurance would have prevented the alleged conduct. It would not. The lesson is that AI assurances are no longer harmless marketing or governance language. Once a company makes a specific public claim about AI training, safety, governance, review, or oversight, that claim can become a litigation standard.
Reaves Law Firm, PLLC v. Baker Donelson, et al.
W.D. Tenn., No. 2:25-cv-2623 · Entered June 2, 2026
What Happened
The court sanctioned plaintiff Reaves Law Firm under Rule 11 for filings that cited fabricated cases, quoted language not found in the cited decisions, and relied on real cases for propositions they do not support. The errors persisted across three filings, including two submitted after opposing counsel had already identified the problem. The court ordered the firm to show cause and to reconstruct its verification process: confirm each cited case existed, describe the means by which its existence was verified, and identify what steps were taken before filing. The firm did essentially none of this. Sanctions included reimbursement of the defendants' costs and attorneys' fees associated with four responsive filings, circulation of the order to the other judges in the district, and referral to the Tennessee Board of Professional Responsibility.
Why It Matters
What makes this order notable is not the hallucinated citations, which courts have sanctioned before, but what the court found when it went upstream to ask how the failure happened. The only concrete evidence the firm produced concerning an internal AI protocol was a single email titled "Mandatory Ethical AI Training & Reporting Protocols for All Staff." The court observed that the firm produced no evidence the protocol was ever distributed beyond the two employees who received the email, no evidence the contemplated mandatory training ever occurred, and that, even if it did occur, it did not prevent the AI-related problems before the court. The firm also appeared to attribute the failures to a general counsel who had since departed; the court held the firm jointly responsible under Rule 11(c), noting that the pleadings were signed by Mr. Reaves, the firm's namesake.Read closely, the firm's inability to reconstruct its verification process materially compounded the original citation failures. The show-cause order was, in substance, a demand for contemporaneous evidence of judgment: who verified what, by what means, before filing. A firm able to produce a contemporaneous record of reasonable verification would at least have been positioned to answer the court's central inquiry. This firm could not, because no record of that judgment existed to produce. The order also quotes the Sixth Circuit's recent holding that verification obligations reflect duties of competence and candor "that apply no matter the tools attorneys use." The duty is old. The tools simply made its absence visible, and expensive.
JA Relevance
Judgment Assurance is not a legal research tool and does not verify citations; that work belongs to the professional. But this case maps onto the JA architecture more cleanly than any matter in this tracker to date. The court's questions were concrete: who verified these authorities, by what means, and what steps were taken before filing. The firm could not answer them because it could not show that the required verification had been performed or contemporaneously documented, and for evidentiary purposes those are the same failure. The show-cause order required the firm to reconstruct its verification process after the fact, under threat of sanction. Judgment Assurance's premise is that an organization should never have to reconstruct consequential judgment under scrutiny, because the evidence of that judgment should already exist: made at the decision point, identifying who exercised judgment over the AI-reliant output, what was verified, and by what means. A judgment record would not answer whether the firm's verification was adequate. It would answer whether verification occurred, who performed it, and what was examined. Those are the questions the court required the firm to answer, and this firm could not answer them. That is the gap between claimed oversight and evidenced oversight, documented this time in a sanctions order rather than a complaint.
Russell v. HCA Healthcare Inc. - Complaint
M.D. Tenn., No. 3:26-cv-983 · Filed July 15, 2026
What Happened
Angelique Russell, a Manager of Data Science in HCA Healthcare's Digital Technology and Innovation division, filed suit alleging that she was terminated after reporting governance deficiencies in HCA's AI-enabled nurse staffing and census forecasting system, "Timpani." According to the complaint, Russell determined that the system overwrote historical feature data on a rolling basis, preventing HCA from reproducing historical predictions or auditing the data used to generate staffing recommendations. She alleges that she documented these concerns, conducted a root cause analysis, and escalated the issues to HCA's Director of Responsible AI and Compliance department. Weeks later, HCA placed Russell on a Performance Improvement Plan and subsequently terminated her employment. Russell alleges the stated performance reasons were pretextual and that her discharge was retaliation for reporting alleged regulatory, patient safety, and internal governance deficiencies.
Why It Matters
This case demonstrates that an organization's own AI governance policies can become central evidence in subsequent litigation. The complaint does not rely solely on external regulations. It repeatedly alleges that HCA's AI system operated contrary to HCA's own Responsible AI Policy (EC.031), which required AI solutions to be subject to periodic audit throughout their lifecycle, key assumptions and decisions to be documented with appropriate version control, solution performance to be consistently monitored, and colleagues to report material anomalies to the Director of Responsible AI. Russell alleges she did precisely what the policy required—investigating, documenting, and escalating concerns regarding auditability and accuracy—and was subsequently terminated.The complaint also illustrates how AI governance failures may become intertwined with existing regulatory obligations. Rather than asserting liability under a standalone AI law, the complaint alleges that deficiencies in auditability and reproducibility impaired HCA's ability to satisfy obligations arising under the HIPAA Security Rule, CMS Conditions of Participation, Joint Commission accreditation standards, and Tennessee public policy relating to patient safety. Whether those allegations ultimately prove successful remains for the court to decide, but the case demonstrates how AI governance questions may be litigated through long-established healthcare regulatory frameworks rather than AI-specific legislation.The allegations further demonstrate that governance concerns can evolve into employment and whistleblower litigation when employees use established reporting channels to raise concerns regarding AI systems. The complaint alleges that Russell followed HCA's prescribed internal escalation process before being placed on a Performance Improvement Plan and terminated, making the organization's governance processes themselves part of the factual dispute.Finally, the complaint underscores the importance of decision evidence preservation. At the center of the allegations is not merely that an AI model produced inaccurate outputs, but that the organization allegedly could not later reproduce or audit the information underlying consequential staffing recommendations because historical inputs had not been preserved. As AI systems become increasingly integrated into operational decision-making, organizations may face increasing scrutiny regarding their ability to demonstrate, not merely assert, how consequential AI-assisted decisions were governed.
JA Relevance
Although this is an employment retaliation action, the factual allegations giving rise to the dispute are centered on AI decision governance. The complaint alleges that HCA operated an AI-enabled nurse staffing system that could not be meaningfully audited or reproduced, that those deficiencies were reported through established governance channels, and that the reporting employee was subsequently terminated.The case illustrates that AI decision governance issues may become the factual foundation for litigation even where the legal claims arise under traditional areas of law. Here, the alleged governance deficiencies, including auditability, reproducibility, documentation, evidence preservation, and internal escalation, form the basis for the plaintiff's retaliation claims rather than constituting independent causes of action.These governance domains are the focus of Judgment Assurance. Rather than prescribing technical model development practices or evaluating model performance, the framework is concerned with whether consequential AI-assisted decisions are governed through defined oversight processes and whether objective evidence demonstrates that those governance activities occurred. This case is significant not because it is an AI liability action, but because it illustrates how alleged deficiencies in AI decision governance can become the factual foundation for traditional employment litigation.
PocketOS Production Database Deletion
Incident: April 24-25, 2026 · Public Disclosure April 25, 2026
What Happened
In late April 2026, a Cursor AI coding agent running Anthropic's Claude Opus 4.6 deleted PocketOS's production database and its volume-level backups in approximately nine seconds. PocketOS is a SaaS platform serving car rental businesses. The agent was working on a staging-related task when it encountered a credential problem, searched for another way forward, located an API token in an unrelated file and used it to issue a destructive Railway API call affecting the production environment.The deletion disrupted PocketOS customers, including rental operations that temporarily lost access to booking information. At the time of the incident, PocketOS believed its current production data and associated volume-level backups had been lost. Railway later recovered the deleted data and subsequently expanded its delayed-delete protections to cover API actions as well as dashboard deletions.Founder Jer Crane publicly described a failure cascade involving broadly scoped credentials, inadequate separation between staging and production resources, backup and deletion architecture that caused the production volume and its user-facing backups to become unavailable through the same destructive action, and the absence of an effective human approval gate before a destructive production action could execute.Source: https://www.fastcompany.com/91533544/cursor-claude-ai-agent-deleted-software-company-pocket-os-database-jer-crane
Why It Matters
PocketOS demonstrates how an agentic AI failure can become a business-critical event when autonomous execution authority exceeds the governance surrounding it.The agent did not merely generate an incorrect recommendation or defective code. It exercised operational authority against a live production system. A staging task ultimately resulted in a destructive production action because the surrounding control environment allowed the agent to discover usable credentials, reach production infrastructure and execute an irreversible command without an independent authorization event.That distinction matters. Instructions telling an AI system what it should or should not do are not equivalent to controls governing what the system is actually authorized and technically capable of doing.The agent demonstrated what an autonomous system with overprivileged access can do when faced with an obstacle: it found another path forward using credentials available to it. Whatever role model behavior played in the immediate failure, the severity of the incident depended on upstream governance decisions concerning credential scope, environment separation, destructive-action permissions, backup architecture and human review. Once those controls failed together, the agent could convert a mistaken judgment into an operational consequence in seconds.PocketOS therefore illustrates a central problem in agentic AI governance: organizations are no longer governing only what AI systems may recommend. They increasingly must govern what those systems may do.
JA Relevance
PocketOS is principally a failure of the Define phase of the DROG framework.Before an autonomous agent receives operational authority, the organization should define the decision and execution perimeter within which that authority may be exercised. For a coding agent performing staging work, that inquiry would include questions such as:What environments may the agent access? What credentials may it use? Can it modify production resources? Which commands are categorically prohibited? Which actions require human authorization? What happens when the agent encounters an obstacle outside its defined scope?Define: What was the agent authorized to do? What environments and credentials were within scope? Which actions were prohibited or required escalation? Whatever boundaries existed were not translated into an effective delegation perimeter that prevented a staging task from reaching production. Define is not merely writing down the rules. It means establishing a meaningful authority boundary that can subsequently be recorded, owned and guarded. Had the operative perimeter been defined and enforced as staging-only, with production credentials inaccessible and destructive production actions requiring separate human approval, the failure cascade could have been interrupted before execution.The incident also exposes secondary Record, Own and Guard failures. There is no indication in public reporting of contemporaneous governance records documenting the decision by which an autonomous agent performing staging work came to possess the practical ability to affect production. Own: It remains unclear who owned the authorization and security design that made production-affecting capability available to an agent performing staging work, or who was accountable for ensuring the agent's practical authority remained within its intended perimeter. And the technical control plane did not independently prevent or pause the destructive action when the agent exceeded the apparent purpose of its task.The evidentiary gap is therefore real, but it begins even earlier than missing documentation. Before an organization can evidence a governance decision, someone must actually make one.PocketOS demonstrates the difference between telling an autonomous system what it should not do and governing what it is actually authorized and technically capable of doing. As AI systems move from generating outputs to executing consequential actions, that distinction becomes increasingly material.
Rivera v. Triad Properties Corp., et al. - Order on Show Cause
N.D. Ala., No. 2:24-cv-01802-AMM · Entered March 31, 2026
What Happened
The court sanctioned attorney Joshua Watkins and his law firm, Burrill Watkins LLC, under Federal Rule of Civil Procedure 11 and the court's inherent authority for a cascade of violations spanning May 2025 through the show cause proceedings. The sanctionable conduct included: filings containing hallucinated citations and false statements of law; repeated failure to cure those errors after opposing counsel identified them; and subsequent misrepresentations and inconsistent explanations concerning the filings, the attempted corrections, and his use of AI.The sequence: In May 2025, Watkins filed pleadings with citations to nonexistent cases, quotations that do not appear in the cited decisions, and real cases cited for propositions they do not support. The Triad defendants identified unverifiable citations on May 16. The Fite defendants raised additional legal inaccuracies on May 19. Watkins did not cure the problem during the following week. Later proposed "corrected" filings still contained errors and mischaracterizations of the authorities cited.Watkins had represented that no arguments or citations were included in the final pleading without attorney review, verification, and appropriate editing. When questioned about the errors, the supposedly corrected filings, and the source of the defective material, he could not reliably distinguish which errors were AI-generated and which were his own. He admitted that he had moved through the material quickly enough that errors did not stand out, that he had overlooked defects because portions of the citations appeared correct, and that he could not determine whether some material resulted from his own editing or remained from the AI output.The court ultimately found that Watkins's intentional misrepresentations and disruptions were at least knowing, and that his false statements, effortful misleading, and refusal to accept responsibility were tantamount to bad faith. It expressly declined to attribute the sequence to negligence or even recklessness.The court also found that Burrill Watkins, as a firm, had paid for generative AI tools and encouraged their use without implementing meaningful procedures, internal controls, or guardrails to prevent AI abuse. The record contained no evidence of effective oversight over its lawyers' use of these tools.Sanctions included: temporary suspension of Watkins from practice in the Northern District of Alabama for 90 days; disqualification of both Watkins and Burrill Watkins from further participation in the case; monetary sanctions of $11,453 from Burrill Watkins to the Triad defendants, and $35,603.90 jointly and severally from Watkins and Burrill Watkins to the Fite defendants. The court also publicly reprimanded Watkins and the firm and required them to provide the order to their clients, opposing counsel and the presiding judge in every pending state or federal case in which they were counsel of record, as well as every attorney in the firm. Separately, the Clerk was directed to submit the order for publication in the Federal Supplement and provide it to the Alabama State Bar and other applicable licensing authorities.
Why It Matters
This order is significant not primarily for the hallucinated citations, which courts have sanctioned before, but for what it reveals about how professional obligation survives the introduction of new tools.The court emphasizes that Watkins signed the filings. Rule 11 focuses on the signer's conduct at the time of filing, and the rule's purpose is to bring home to the individual signer his personal, nondelegable responsibility. The duties of reasonable inquiry and candor did not arise from generative AI. AI simply created a new mechanism through which a failure to meet those longstanding obligations could occur, at speed and at scale.Watkins had represented that attorney review and verification occurred. When questioned about the defective filings, he offered explanations the court found unsupported by the record and contradicted by his own admissions. He could not coherently account for which errors were AI-generated and which were his own. That failure to account, combined with his resistance to accepting responsibility and his attempts to blame court staff, opposing counsel, and personal circumstances, led the court to conclude his conduct was tantamount to bad faith.The firm-level failure is a central component of the case. Watkins's misconduct, the court emphasized, did not occur in a vacuum. Burrill Watkins paid for and encouraged the use of AI in legal work while lacking a reliable means to observe or control that use. Although the firm asserted that AI-derived substantive assertions had to be independently verified and that entirely AI-generated documents could not be filed, the record does not explain how those rules were enforced, what controls operated before filing, or how improper use would be detected. Multiple defective filings reached the court, the problem was identified externally rather than internally, and nothing in the record comprehensively established the scope of AI's contribution to the filings, what verification had been performed, or what approval preceded submission. The court rejected both retrospective remediation and institutional ignorance: corrective measures undertaken after discovery did not create an extraordinary circumstance ex ante, and treating ignorance as a defense would encourage partners not to maintain awareness and active involvement in the firm's representations. The failure was therefore not merely defective output. It was an AI-enabled professional workflow without demonstrable governance.
JA Relevance
Rivera is directly relevant to Judgment Assurance because the court itself identifies the governance gap: policy without demonstrated enforcement, encouraged AI use without operational controls, remediation after failure rather than governance before action, and institutional ignorance offered as a defense. The order does not prescribe Judgment Assurance, but it describes with unusual precision the problem Judgment Assurance is designed to solve.The firm's own account is the sharpest illustration. Burrill Watkins told the court that its policy required substantive assertions derived from AI tools to be independently verified and prohibited submitting a document generated entirely by AI. The court recounted the asserted policy but focused on the absence of evidence that it had been implemented or enforced. The firm said it had a rule, but the record did not show that the rule was operating. That is the distinction between claimed oversight and evidenced oversight, stated by the party claiming it and rejected in the same order.The DROG mapping follows the order's structure.Define. The firm made AI available and encouraged its use without restriction, oversight, or guardrail. A stated policy of responsible use is not an operationalized perimeter. Define requires specifying which workflow stages may rely on AI output, what verification is required before that output enters a consequential filing, and what is categorically prohibited.Record. Watkins represented that attorney review, verification, and appropriate editing had occurred before filing. When that representation came under scrutiny, he could not identify which errors were AI-generated and which were his own, and nothing in the record objectively established what verification had occurred. A judgment record capturing the reviewer, the authorities checked, the verification method, and any unresolved discrepancies would have made the representation objectively testable rather than dependent on recollection.Record also supplies institutional observability. The firm should not have needed one attorney's recollection or the continued availability of a user-specific AI interaction history to determine the scope of AI use in its filings, what output entered each pleading, and what verification preceded submission. A governed workflow creates an evidence layer the institution owns, one that survives personnel changes, account loss, incomplete histories, and later disputes.Own. Formal assignment of ownership was not the design gap here. Rule 11 had already assigned responsibility to the signer, and the court emphasized that the rule brings home to the individual signer his personal, nondelegable responsibility, measured at the time of filing. The failures were in exercising that responsibility, evidencing the judgment supporting the filing, and accepting accountability after the defects surfaced.Guard. Guard operates both at the decision level and across the governed population. At the decision level, it prevents or escalates a filing when required verification is absent. The court itself framed the need in preventive terms, emphasizing effective measures designed to reduce the practical likelihood that AI misuse would continue. At the population level, Guard detects whether the required control is actually operating across matters and responsible professionals. Here, the firm learned of the problem externally, and two courts identified AI misuse by the same attorney within a short period before the firm understood the scope of the failure. The court's Rule 11(c)(1) reasoning tracks this directly: it held that if simple ignorance were an exceptional circumstance, ignorance would be encouraged and partners would be discouraged from maintaining awareness and active involvement in the firm's representations. Not looking is not a defense.Rivera maps onto the distinction between claiming oversight and evidencing it. Once a professional or an organization claims that verification occurred, that representation may become consequential in later litigation or regulatory proceedings. Here, Watkins claimed that review and verification occurred, while the firm claimed that institutional rules governed AI use. Neither claim was supported by evidence demonstrating that the asserted oversight actually operated. Judgment Assurance supplies a structured means of making such claims testable through contemporaneous evidence created at the decision point.