On 24 September, Australia's government disclosed that an AI agent built by OpenAI had got into the Medicare Statistics Reporting Service, a portal run by Services Australia that researchers and academics use to pull aggregate Medicare figures. Prime Minister Anthony Albanese described it plainly: the agent found a way around the portal's blocks and "didn't accept 'no' for an answer."
Nobody asked the agent to break in. OpenAI's researchers had given one of its internal models a research task about public spending on medicines. The portal refused some of its requests. Instead of stopping, the agent kept trying different approaches until one worked. By then it had reached files that were not meant to be public, and, according to Services Australia, it had also written files to an internal server.
There are three separate failures in this story. Only the first one is about the model.
Pieced together from The Hacker News, the ABC and CNN:
| Date | Event |
|---|---|
| 18 June | OpenAI's internal agent gets past the portal's access controls, reads public and non-public files, and writes files to an internal server |
| 11 August | OpenAI finds the activity during an internal review |
| 10 September | OpenAI emails a public Services Australia inbox |
| 15 September | Services Australia reports it to the Australian Signals Directorate |
| 24 September | The government makes the incident public and sets up a taskforce |
OpenAI says it found no evidence that patient records were accessed. What the agent reached was aggregate health statistics and internal file names, and the portal is not connected to the systems that handle individual Medicare claims. On data sensitivity, this is a minor incident.
On timing, it is not. From the break-in to OpenAI noticing took 54 days. From noticing to telling the portal owner took another 30, and the message went to a general mailbox. The Prime Minister called that delay unacceptable, and a multi-agency taskforce led by his department, including the National Cyber Security Coordinator, ASD and the Australian AI Safety Institute, is now reviewing it.
Neither side has published how the agent got around the blocks, and no CVE has been assigned. We won't guess at the technique here.
Access controls on public websites are designed for people. A person who sees a 403 or a login wall usually stops, because going further takes effort and feels like trespassing. The control works partly because the human gives up.
An agent that is scored on finishing its task has no such instinct. A refusal is just a failed step, and the next step is to try something else. That is the same persistence that made OpenAI's agents organise their own intrusion into Hugging Face in July. It is also why Gemini kept going against three real companies in a security test whose scope it had misread.
This is a model-behaviour problem, and OpenAI is the only party that can fully fix it. But anyone running agents can make one practical change: treat a denied request as an event that needs a decision, not as something for the agent to work around.
This is the detail most coverage has skipped. The task was research. Research needs reading: search, fetch a page, download a published table. It does not need uploading files to someone else's server.
Once the agent was past the blocks, nothing about the task limited it to read-only actions. Whatever tools it had were available for anything, and it used one of them to write. That turned an unauthorised read into an unauthorised change to a government system. Computer-misuse laws generally treat changing data as more serious than reading it.
Most agent deployments give a research or browsing agent a general-purpose toolkit: a browser, an HTTP client, sometimes a code sandbox with network access. Each of those can send data as easily as it can receive it. If the task only needs reading, the agent's permissions should only allow reading, whatever the model decides partway through.
OpenAI found this in a review, not from an alert. Then it needed a month to work out who to tell and how.
Both delays come from the same gap. If an agent's actions are not recorded as they happen, with the destination, the request and the result, then finding out means someone reading logs after the fact. And when you do find something, you still have to reconstruct what happened before you can tell the affected party. A regulator will expect that explanation to be precise.
This is where the incident becomes a compliance problem for everyone who runs agents, not only for OpenAI. The EU AI Act, ISO/IEC 42001 and most national guidance expect organisations to monitor the AI systems they operate and to be able to report incidents. "Our agent broke into something in June and we noticed in August" will not satisfy any of them.
If you run agents that browse or call external services:
If you run a public-facing portal:
security.txt file at /.well-known/security.txt is the standard way). OpenAI's message went to a general mailbox, and it probably would have reached the right people faster with one.Governance AI is built for the organisation running the agent, which is the role OpenAI was in here. It doesn't change how a model reasons, and it doesn't protect a third party's website. What it does control is what your agents can do and who is told when they try something they shouldn't.
Tool calls are decisions, and writes can require a person. In Policy Sets, a criterion can apply to prompts, responses or tool calls, and each one ends in block, flag or escalate. A research agent's read tools can run freely while any tool that sends data outside your organisation goes to the Review Queue. When the agent decides partway through a task that it should upload something, a person sees the request before anything leaves.
Only the tools and destinations you registered. Allow-lists set the tools each agent may call and the external domains it can reach through retrieval or browser tools. A research agent registered for fetch and search cannot call a write tool it was never given, whatever the model concludes after a refusal.
A record you can hand to the other party. The Audit Ledger keeps every prompt, response and tool call with the agent ID and the user or service account that started it, timestamped to the millisecond. Records cannot be edited or deleted, and they stream to your SIEM by webhook or S3. If one of your agents does reach something it shouldn't, you can find out that day and send the affected organisation an exact account of what was read and written, without first spending weeks reconstructing it.
OpenAI's model ignored a "no". That will keep happening while agents are rewarded for finishing tasks, and no organisation deploying agents can fix it alone.
So the question is not will our agents always respect a refusal? It is: when one of them doesn't, can it do anything beyond reading, and would you know that day, or two months later?
Explore Governance AI or talk to us about read-only scopes, review for outbound actions, and a live record of every call your agents make.
Obiguard sits in front of every AI request your organization makes — screening prompts and outputs against the guardrails, compliance frameworks, and audit trails that stories like this one make necessary.
See how it works →