Frontier labs do not usually coordinate their releases. In the first two days of September 2026 they may as well have.
Over September 1 and 2, OpenAI, Google and Anthropic each announced a model with frontier-grade offensive security capability, and each announced an access programme to decide who is allowed to use it. The Hacker News covered all three together on September 2, which is the right way to read them — not as three product launches, but as one industry-wide admission arriving in triplicate.
The admission is this: autonomous discovery of previously unknown vulnerabilities in hardened systems is no longer a research result. It is a shipping capability with a waiting list.
OpenAI confirmed that its forthcoming Astra model meets the Critical cybersecurity capability threshold under its Preparedness Framework — the first model in the company's history to do so. The threshold is defined as a system that can "identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks." This was not a surprise landing; OpenAI had flagged in August that critical capability "cannot be ruled out," and paused parts of Astra's development while it strengthened safeguards.
The numbers attached to the confirmation are worth sitting with. Astra scores 100% on ExploitBench. It declines 91.5% of jailbreak attempts, against 59% for GPT-5.6 Sol. In evaluation it found previously unknown flaws and assembled them into working chains — including complete browser sandbox escapes and privilege escalation chains. The narrow slice of capability that triggered the Critical rating is restricted to vetted defenders through Daybreak Blue, initially a small alpha group. Every Daybreak account has been required to use a hardware security key since September 1.
Google introduced Gemini 3.8 Flash Cyber and the Fairwind Program, which gives early access to more than 650 partners — Google Cloud customers, government agencies and security vendors including CrowdStrike, Datadog, Menlo Security, Palo Alto Networks and Snowflake. Google reports frontier-level performance on the CyberGym benchmark and a success rate above 70% on an internal benchmark spanning 20 languages. Two partner data points stand out: Chrome's security team got 2.6× more correct patches than from much larger commercial models, and Wiz recorded 7.5–9.7% higher recall on its internal pentest benchmark at 2.3–5.2× lower cost.
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 — the same underlying weights, split by safeguard configuration. Fable 5.1 is generally available. Mythos 5.1 is restricted to vetted organisations through a Cyber Verification Program, currently a limited set of US organisations. Alongside it came Enterprise Frontier Safeguards, which pairs zero-data-retention-equivalent privacy with misuse detection by keeping data in infrastructure the customer controls, rolling out in phases later this autumn. Anthropic also shipped sandbox-escape detection classifiers and reward-mechanism changes — a direct response to the containment failures we covered when 1,200 agents coordinated on an unsanctioned message board.
Two weeks ago we wrote about five US federal agencies warning that AI-generated Python was being used against Siemens S7 PLCs with decade-old exposures. The argument there was that the barrier protecting those systems was never the vulnerability — it was the scarcity of people who could weaponise it, and that scarcity had quietly ended.
These three announcements are the same fact stated from the supply side, deliberately, by the suppliers. The labs are not disputing that the expertise is now synthesisable. They are building turnstiles in front of it.
The turnstiles are real and they are better than nothing. But look carefully at what they certify. A vetting programme confirms that an organisation is a legitimate defender. It says nothing about:
Every one of those is your control, not the lab's. The vetting boundary sits at their door. The exposure sits behind yours.
Here is the part that gets missed in coverage framed as defenders get better tools.
Defensive work with a cyber-capable model means feeding it your most sensitive material. Source code. Unpatched findings with the fix not yet shipped. Network topology. Incident timelines. Credentials rotated mid-response. That corpus is, almost exactly, the target package an attacker would spend months assembling.
Now consider the distribution. Daybreak Blue starts as a small alpha. Mythos 5.1 is limited to a vetted set of US organisations. Fairwind has 650 partners, which sounds large until you weigh it against the number of organisations that run a SOC.
Most security teams will not be admitted to any of these programmes this year. They will still do the work — because the productivity gap is real, and because the generally-available tiers, Fable 5.1 and Gemini 3.8 Flash and GPT-5.6 Sol, are extremely capable at code review and log analysis and malware triage even without the restricted slice.
They will do it in whatever window is open. Which, for most organisations, is a personal account.
That is shadow AI arriving in the one department where it costs the most. IBM's 2026 figures put shadow AI in 43% of studied incidents, roughly double the prior year, and found only a minority of organisations have any technical control preventing uploads to public AI tools at all. Apply that base rate to the security function specifically and the exposure is not "an employee pasted a customer record." It is your unremediated vulnerability inventory sitting in a consumer chat history that your DLP never saw, because a prompt looks like ordinary encrypted web traffic.
If question 4 has no answer, questions 1 through 3 are academic.
The gap this story opens is not an offensive-capability problem you can buy your way out of. It is a plain governance problem: capable models now exist, your people will use them, and the only variable you control is the surface they use them through.
Obichat is that surface. It is a governed AI workspace built on LibreChat — the familiar interface people actually want — with the controls that make it safe to say yes.
Per-workspace model allow-lists. Each team gets its own workspace with its own connected providers, and IT decides which models that workspace may reach. A security team's workspace can be bound to exactly the vetted model you are approved for — no rogue API keys, no unmanaged endpoints, and no ambiguity about which tier of a split-safeguard model pair like Fable/Mythos is actually in use.
Policy enforced before the prompt leaves your network. Every message passes through Obiguard's inspection layer under the policy set assigned to that workspace: PII and PCI redaction, keyword and regex block-lists, jailbreak and prompt-injection detection. For a SOC workspace, the block-list is where "never send raw credentials or customer identifiers to a model" stops being a training slide and becomes a control.
An exportable record per workspace. Conversation and event history, exportable to CSV, with framework tagging for NIST AI RMF and ISO/IEC 42001. That is question 4, answered as a query.
It fits the identity setup you already have — SAML 2.0 and OIDC sign-in, admin-gated workspace settings, workspaces scoped private or team-wide, and deployment in your own AWS, Azure or GCP tenant if the data cannot leave your perimeter. Paired with firewall-level blocking of consumer AI endpoints, Obichat becomes the only on-ramp rather than the preferred one.
Two adjacent halves matter here. For the machine side of the estate — agents and services rather than people — Governance AI applies the same logic through allow-lists that bind credentials to specific model IDs, tools, domains and invoking identities, with an immutable audit ledger streaming to your SIEM. That is the control we argued for when an AI coding agent did the lateral movement in a live ransomware intrusion. And on the receiving end of all this new capability, Obiguard SOC is where the findings land — CVE Radar on every push and daily, cross-referenced against CISA KEV and FIRST.org EPSS.
The comforting reading of this week is that defenders got a big upgrade. Chrome's 2.6× patch improvement is real. Wiz's recall numbers are real.
The uncomfortable reading is that all three labs simultaneously decided this capability needs a gate — and a gate is an admission that the thing behind it is dangerous enough to warrant one. Gates leak. Weights get replicated, open models close the distance a few months later, and every vetted account is a credential someone can phish. OpenAI mandating hardware keys across Daybreak from September 1 is not a formality; it is the lab telling you exactly how it expects the gate to fail.
So the durable control was never going to be the gate. It is knowing, on your side of it, which models your people can reach, what they are permitted to send, and what was actually sent.
Right now, for most organisations, the honest answer to all three is the same: nobody has looked.
Explore Obichat or talk to us about what your security team sent to an AI model last quarter — and whether you could prove it.
Obiguard sits in front of every AI request your organization makes — screening prompts and outputs against the guardrails, compliance frameworks, and audit trails that stories like this one make necessary.
See how it works →