Commentary

Agentic AI Has No Paper Trail

The missing logs

6 min read

On October 24, 2025, NIST released the second public draft of SP 800-53 Concept Paper for Control Overlays for Securing AI Systems, the agency's most detailed attempt yet to translate existing federal cybersecurity controls into something an autonomous agent operator can actually implement. Buried in the overlay scope for generative and agentic systems is a control family the document treats as load bearing: AU, the audit and accountability family. The overlay specifies that agentic systems must log tool invocations, model outputs, prompt chains, retrieval calls, and the identity of the principal on whose behalf the agent is acting. The draft is open for comment. Almost no production agent stack on the market today satisfies it.

That gap is the story.

The failure mode at the center of autonomous agent accountability is not hallucination, jailbreak, or prompt injection, though each gets more press. The failure mode is the absence of a record. When an agent books a flight, files a return, sends an email under a delegated credential, or executes a trade through a connected brokerage API, the question that determines whether anyone can be held responsible is whether a tamper evident log exists showing what the agent did, what it was instructed to do, and what data it consulted before acting. In most current deployments, the answer is partial at best. Vendors log enough to debug their own product. They do not log enough to let a regulator, a plaintiff, or an internal auditor reconstruct the decision.

Consider the receipt trail from the past year. On May 10, 2025, OpenAI disclosed that a third party analytics provider, Mixpanel, had suffered a security incident exposing limited API user metadata, including names, email addresses, and approximate location data. The notice identified the data category and gave a remediation timeline. It did not provide per tenant logs sufficient for affected enterprises to determine which of their own users' agent sessions touched the compromised pipeline. Customers were told, in effect, to assume exposure and rotate. That is not accountability. That is notification theater built on the assumption that nobody downstream keeps records granular enough to ask harder questions.

Anthropic's October 22, 2024 announcement, "Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku," is more candid and, for that reason, more revealing. The company acknowledges that when an agent operates a browser or filesystem on a user's behalf, the security boundary collapses into the user's own session. Anthropic logs API calls. Anthropic does not, and structurally cannot, log what happens inside the user's operating system once the agent has the keys. The post recommends sandboxing, human confirmation for sensitive actions, and avoiding exposure of sensitive credentials. Each recommendation is sound. None produces the audit trail NIST is asking for. The trail has to be built by the deployer, and the deployer is usually a small firm that bought an agent platform precisely because it did not want to build infrastructure.

This is where the Transfer Ratio shows up. The Transfer Ratio is the share of a product's price that funds the accountability surface, the logging, retention, and forensic readiness a downstream auditor would need, as opposed to the share funding model access, orchestration, and integration. Pull published October 2025 list pricing from three platforms: OpenAI ChatGPT Enterprise at roughly $60 per seat per month, Anthropic Claude Enterprise at roughly $75 per seat per month, and Microsoft Copilot Studio at $200 per tenant per month plus message pack consumption. Each SKU bundles standard retention of 30 to 90 days, no tamper evident export, and no chain of custody attestation in the base price. Forensic grade logging, where offered at all, sits behind separate compliance add ons or enterprise data residency tiers priced in the low single digit dollars per seat. The audit slice lands near three cents on the dollar. The vendors are not hiding this. The line items are right there in the SKUs. What the vendors are doing is selling a product whose accountability surface depends on a component the buyer is not buying and does not know to ask for. When something goes wrong, the buyer discovers that the relevant logs either do not exist, exist for thirty days, or exist in a format that fails the authentication standard set by Federal Rule of Evidence 901 and the self authentication path for machine generated records under Rule 902(13).

Call this what it is. It is a cost shift. The expense of producing a defensible record of automated decisions has been moved off the vendor's balance sheet and onto a future plaintiff, a future regulator, or a future bankruptcy estate. Theft by Another Name, the framework that names a transfer of value from a diffuse group to a concentrated one through a system whose rules were written by the concentrated group, applies in the narrow sense it is built for. End users and counterparties exposed to agent error are subsidizing the platform vendors who avoid the cost of building real audit infrastructure. The system is visible. No malice has to be assumed.

Responsibility for the absence sits in three places, and naming them matters.

First, the platform vendors themselves. OpenAI, Anthropic, Google, and Microsoft have the engineering capacity to ship tamper evident logging by default. They have chosen not to, because default logging raises storage costs, complicates the privacy story, and surfaces failure modes the marketing copy prefers to elide. Each firm has published responsible scaling or trust and safety documentation. None of those documents commits to a logging standard a third party could verify.

Second, the federal agencies that were supposed to set the floor. The October 2025 NIST overlay is welcome and overdue. It is also a draft, it is voluntary, and it arrives nearly three years after ChatGPT shipped and roughly eighteen months into the agentic deployment wave. EO 14110, the Biden executive order on AI, was rescinded by President Trump on January 20, 2025 through Executive Order 14148. The replacement, "America's AI Action Plan," released by the White House on July 23, 2025, emphasizes deregulation and federal procurement leverage rather than enforceable logging standards. The result is a federal posture in which the technical document with the right answer, the SP 800-53 overlay, has no statutory teeth, and the policy document with reach points away from the problem.

Third, the procurement officers and general counsels at the firms deploying these agents. The contracts being signed in 2025 routinely accept vendor side logging as sufficient. They do not require log export, do not specify retention floors, and do not contemplate the chain of custody questions a regulator will ask in 2027. This is a failure of legal imagination, and it is the failure most easily corrected. A procurement template that requires SP 800-53 AU-2, AU-3, AU-9, and AU-12 compliance, with logs exportable to customer controlled storage, would change vendor behavior inside two quarters. No statute is required. Buyers have the leverage and are not using it.

The pattern repeats across every wave of computing infrastructure that preceded this one. Email, cloud, mobile, each shipped without adequate logging, each was retrofitted under regulatory pressure after a sufficient number of public failures, and each retrofit cost more than building the capability in the first place. Autonomous agents are early in that arc. The NIST draft is the first real attempt to short circuit the cycle.

Whether it works depends on whether the document acquires force, through procurement language, through state attorneys general, through the first major piece of agent driven litigation. Until then the logs are not being kept, the receipts are not being filed, and the firms most responsible for the absence are the firms that will, in the first wave of accountability cases, claim the records simply are not available.

They will be telling the truth. That is the problem.

What say you?


Originally published at https://henrygoodstone.com/commentary/agentic-ai-has-no-paper-trail/

Get every essay by email. Subscribe free on Substack.