|
||
|
||
The OpenAI, Anthropic, Alibaba, and now AISI containment failures are not all the same. And the difference makes a strong case for a common-sense governance mechanism that we can implement today. The difference isn’t in what the models did. It’s about who was watching. It was standard network infrastructure, not specialized frontier AI-tooling, that meant the difference between the longest discovery gap of roughly three months and the shortest, which caught the offending model in the act. The argument I am making here is that we have the shape of accountability infrastructure for AI now. We can cheaply and rapidly adapt existing mechanisms to AI inference traffic and make a major advance in the governance capability available to private stakeholders and governments alike.
In a now internet-infamous incident at Alibaba that was quietly included in a ninety-author paper, published on arXiv back in the AI stone age of just over seven months ago, a 30-billion-parameter model earned the “rogue” moniker more than any of the recent follow-up acts. No doomer could top the drama provided by the paper’s own authors. They say they observed the AI autonomously repurposing GPU capacity for cryptocurrency mining and “introducing clear legal and reputational exposure.”
That’s the understandably famous shock headline. But a quick scan back to its preceding paragraph reveals a win every CISO and SOC2 auditor should still be applauding today. “Our first signal came not from training curves but from production-grade security telemetry.” The champion elite blue-team countermeasure? A firewall. “Alibaba Cloud’s managed firewall flagged a burst of security-policy violations…”
On July 30, Anthropic offered its version as part of a series of “hold my beer” frontier AI lab containment-fail disclosures. They open by admitting that they hadn’t chosen to review their logs until they saw OpenAI’s kerfuffle, and someone had the bright idea that it may have already happened to them too, without them knowing. Sure enough, not once or twice, but thrice, their models under evaluation compromised the production systems of three real, third-party, minding-their-own-business companies. The earliest incident occurring in April. They found out in late July. The quiet revelation in this account? The transcripts and network logs were on their own servers, and for at least three months they had no plan to look at them.
In every case, the one surface where the problem was always visible? Ordinary network infrastructure.
A firewall caught Alibaba’s crypto mining spree because that kind of misaligned act comes in hot and heavy on a network activity graph. The harms that deserve society’s utmost attention walk straight through the current network corridors undetected.
When a child makes a chatbot their BFF, and it tells them they’re right to confide in them rather than their parents, there’s no spike on a chart. No chaperone on the wire. Today’s network defenses are built for intrusion and data exfiltration attacks, not social or psychological ones.
I have already made the case that AI communication is not exempt from traveling the same routes that we’ve established for any other internet enabled technology, and it leaves footprints. The problem is that none of the current network management regimes were designed to detect or respond to inference-based threats, and no third party in today’s inference transactions bears a duty to the person on the consumer side.
Although that is true today, the demand for legislative action is growing, and AI output prohibitions have already been authored in at least five US jurisdictions. But in each case, the provider is responsible for checking itself. Washington’s HB 2225 forbids manipulative engagement techniques on minors; the operator evaluates whether its own model complies. Four other states passed versions of this, and a federal bill is pending. Lawmakers have carved out ban lists, but no one has yet been given any means to check them before delivery.
Brandie Nonnecke, of Americans for Responsible Innovation, called on July 28 for the US Congress to institute a “chatbot duty of care,” under which developers would “continuously monitor for emerging harms.” I certainly agree with the duty. But monitoring is the part the top providers just failed at, spectacularly. These companies at the bleeding edge of human technology were capable of greater care. The incentives driving their choices to only invoke basic measures after harms were exposed are the same incentives which exclude them from being trusted to stop the harmful outputs already reaching the children.
The provider’s conflicting incentive doesn’t mean that a new bureaucratic regime with a slow reporting mechanism is needed. The proof lies in the frontier labs’ own disclosures. Delivery-boundary controls already exist today. Anthropic’s own disclosure included an admission that “the safeguards deployed on our generally available models would have blocked the behaviors identified.” The evaluation happening at the boundary also provides a record. This means interception and evidence are one thing, not two.
The pattern that turned that fact into a monitoring protocol has already been deployed in the case of email. The standard network rules which the industry itself imposed provide for sender provenance verification, aggregate reputation reporting, and delivery policy enforcement. These capabilities enable spam filtering, which acts not only on end-user delivery, but on aggregate provider reputation. These are basic visibility and intervention capabilities which are glaringly, and at this point inexcusably, missing from the standards imposed on AI inference transactions. It is clearly unreasonable to suggest that frontier AI labs, or any party capable of setting up an AI inference serving endpoint, are incapable of implementing a network protocol they already have for their own email service.
Therefore, I argue it is irresponsible not to require at least such a basic protocol for AI, enabling a duty of care to consumers. A policy which achieves this does not need a heavy-handed, broad mandate across the entire industry, which would be difficult to enforce. It could employ the same tactic that yielded near universal adoption in email protocols. A handful of large mail providers required specific protocols for anyone delivering to their users. The rest of the industry followed because the alternative was not reaching anyone. Government buying power can play the role the large providers played. Control access to a significant portion of the market, and adoption everywhere else becomes unavoidable. The two surfaces where rules are already being demanded are government procurement contracts for AI inference, and the operation of AI inference services which can be accessed by minors. The first is a purchasing condition. The second is a condition of market access. Anyone who wants to sell AI inference to the government, or operate an AI service children can reach, delivers it through a protocol that meets the minimum capabilities to enforce a duty of care.
A specification for this already exists. I wrote it, and submitted it to the IETF as an Internet-Draft, with an open-source implementation to prove it runs. It is not a research proposal. The parts are ordinary, and the reason I can say it is buildable and cheap is that I built it.
There is a stronger version of this ask, and it is worth mentioning so nobody thinks it was overlooked. The check that stops a harmful output could be run by a party that answers to the user instead of the provider. That is a duty of loyalty, and it’s different from a duty of care. A duty of care requires companies not to hurt you. A duty of loyalty requires someone to be on your side. US lawmakers are nowhere near that conversation, and pushing it holds back the smaller thing that is actually attainable now. This layer has to exist before anyone can argue about who holds it.
Pushed by consumer complaints, the email industry decided spam was unacceptable and built the answer into the plumbing. Nobody had to invent an exotic solution. They just decided, and then required it. Anthropic’s transcripts sat on its own servers for three months. Alibaba’s firewall caught a model stealing GPUs in the time it takes to file an alert. The problem is not the lack of technology. It’s the lack of an incentive to guard the boundary between a chatbot and a child. Requiring a common-sense protocol between AI and children supplies the incentive to enable practical AI governance today.
Disclosure: The author develops education technology and is the author of an Internet-Draft submitted to the IETF specifying a provenance, verification, and reputation standard for AI inference, along with an open-source reference implementation. The author therefore has a direct interest in the kind of infrastructure described here.
Sponsored byVerisign
Sponsored byCSC
Sponsored byIPv4.Global
Sponsored byDNIB.com
Sponsored byWhoisXML API
Sponsored byRadix
Sponsored byVerisign