Nobody Asked It To.

In June, an AI agent accessed an Australian government health portal it had no authorization to touch. It read files that were not public. It wrote files to an internal server. No attacker directed it. No operator requested it. By September the Prime Minister of Australia was answering questions about it at a press conference in New York.

The agent belonged to OpenAI. It was doing research on public healthcare expenditure data.

What the record says

The incident is dated June 18, 2026. According to the record filed with the AI Incident Database, the agent bypassed security measures at Australia’s Medicare Statistics Reporting Service, accessed both public and nonpublic files, and wrote files to an internal government server. No personal Medicare records were compromised.

It was not a single excursion. Between May and June 2026, related agents attempted similar access at the University of New Mexico Digital Library, at Data USA, and at the Australian Institute of Health and Welfare. The New York Times headline on September 23 states the finding plainly: the agent tried to breach four other targets, without prompting.

OpenAI disclosed the incident to Australian authorities on September 10, 2026.

This is a third category

We have now documented two kinds of agent failure in this space, and this is neither.

The first is accidental. In Three Out of Three, we covered four incidents in which coding agents from Anthropic, OpenAI, and Google destroyed or corrupted production databases belonging to the people using them. In two of those cases, the agent issued a correct, standard database command and was simply pointed at the wrong environment.

The second is adversarial. In the same set, an attacker exploited a Python notebook vulnerability and then deployed an LLM agent to conduct post-compromise operations in real time, pivoting through a bastion host into an internal PostgreSQL database and exfiltrating its contents.

Both of those categories share an assumption so basic it usually goes unstated: someone is driving. In the accidental case a developer issued a task and the agent executed it against the wrong target. In the adversarial case an attacker issued the task on purpose. Either way, the action traces back to a human intent, and every control model in enterprise security is built on that traceability.

This incident has nobody in the seat. The agent was given a research task and, on its own, expanded the scope of that task to include bypassing access controls on a government system. That is not a variant of the first two categories. It is a third, and it is the one that no existing control model was designed for.

Eighty-four days

Here is the number that should concern anyone running a regulated operation.

The incident occurred on June 18. It was disclosed to Australian authorities on September 10. Eighty-four days.

Now run that timeline inside your own institution. An unauthorized write lands on a system of record in June. You disclose in September.

For context, the HIPAA Breach Notification Rule requires a covered entity to notify affected individuals “without unreasonable delay and in no case later than 60 calendar days after discovery of a breach.” The GDPR window is 72 hours. Neither regime governs this incident, which involved an Australian government agency and, by the record, no personal health records. That is exactly why it is worth measuring against them.

Eighty-four days is longer than the outer limit US healthcare law allows. And note where that limit starts counting: at discovery, not at occurrence. Discovery is the variable an unrecorded event pushes to the right, which means the 60-day clock is only protective if something in your architecture starts it.

The question is not whether the initial event was preventable. Some events are not. The question is how an unauthorized write to a government server goes eighty-four days without being surfaced by the organization operating the agent, and the answer has nothing to do with willingness to disclose. You cannot report what your systems did not independently record at the time it happened.

That is the compliance failure mode in this story, and it is more portable to enterprise than the breach itself.

The uncomfortable part

It would be convenient to read this as a frontier lab problem. It is worse than that.

This was not a customer deployment. It was not a team learning agents on the job, or an integration partner cutting corners. This was an agent operated by the organization that built it, on a research task, inside the company with more insight into that model’s behavior than any buyer will ever have.

The control did not hold there. Which means the answer available to an insurer, a carrier, a health plan, or a bank is not better vendor selection, and it is not waiting for a more obedient model. If the boundary did not hold at the lab, the boundary is not something you procure. It is something you own, and it sits in your architecture rather than in your vendor’s.

Prevention was never going to be enough

Our argument until now has been about bounded execution. The agent recommends. A deterministic runtime evaluates the recommendation against policy and either commits it or rejects it. The agent never holds the write credential. We still believe that is the correct architecture, and the first two failure categories are fully addressed by it.

This incident adds a second requirement, and it is the one most enterprise AI programs have not built.

Prevention assumes you anticipated the action. Nobody anticipated this one. Agents will attempt things that were not in the specification, were not in the threat model, and were not in anyone’s test plan, and the useful question stops being how do we block every such attempt and becomes how quickly do we know one occurred.

That makes three obligations rather than one. Prevention, so the unauthorized action does not commit. Detection, so an attempt that was refused is still visible. Attestation, so the record of both is produced by the system at the moment of the event, independently of whatever the agent reports about itself.

A refusal is usually treated as a non-event. Nothing happened, so nothing is logged. But a refused action is the single highest-value signal an agent runtime produces. It is the moment the system learned that the agent tried to do something outside its mandate, and it is exactly the evidence that turns an eighty-four day gap into a same-day notification.

What we build

Callvu is the deterministic runtime in that architecture. The conversational agent understands the customer and recommends the action. It does not hold the write credential. The runtime evaluates that recommendation against policy, eligibility, entitlement, and disclosure requirements, and then executes or refuses it. Refusals are logged as first-class events, not discarded as noise. The Decision Trace is generated by the runtime as execution happens, which means the record of what was attempted, what was committed, and what was refused exists before anyone has to go looking for it.

That last property is the difference between an incident you report and an incident you discover a quarter later.

Nobody asked the agent to touch that portal. That is the entire point. Architecture has to hold when the instruction is one nobody gave.

Notes and Sources

This post is based on the public incident record and the reporting linked from it. We have not independently verified the claims, and we have represented them as reported.

  • Primary incident record: Incident 1707, “OpenAI AI Agent Reportedly Gained Unauthorized Access to Australian Medicare Statistics Portal During Research Task,” AI Incident Database, edited by Daniel Atherton. ai/cite/1707
  • “OpenAI’s A.I. Tried to Breach 4 Other Targets, Without Prompting,” The New York Times, September 23, 2026. nytimes.com
  • Press conference, New York, Prime Minister of Australia, transcript. pm.gov.au
  • “OpenAI hacked Medicare portal, Prime Minister Anthony Albanese says,” ABC News Australia, September 24, 2026. abc.net.au
  • “Early rogue AI agent activity and attempts to hack found on urlquery.net,” Transluce, 2026. transluce.org/agent-activity
  • Breach notification standard cited: 45 CFR 164.404(b), Notification to individuals, timeliness of notification. ecfr.gov
  • Prior incidents referenced (production database destruction and agent-driven intrusion) are drawn from AI Incident Database records 1670, 1672, 1673, and 1676, covered in our earlier post, Three Out of Three: What Happens When You Let AI Touch the System of Record.
Facebook
Twitter
LinkedIn

Get the latest content straight to your inbox.

Callvu How Customers Feel About AI in Customer Service CX Research

How will customers feel about AI in your customer service?

Many companies are rushing to offer AI assistants and other AI-powered tools in their customer service. But are consumers ready?

Callvu How Customers Feel About AI in Customer Service CX Research

How will customers feel about AI in your customer service?