Press release

Obiguard launches AI red-teaming and evaluations, built on Project Moonshot, in Governance AI

Attack your AI agents with automated red-team tests and grade them against benchmarks before release, with the evidence behind every result kept for review.

KUALA LUMPUR — Obiguard, the AI risk governance company, today made automated AI red-teaming and benchmark evaluations available in Governance AI. Built on Project Moonshot, the open-source toolkit from the AI Verify Foundation, they let teams test how an AI agent behaves under attack, and how it scores on benchmark tests, before it reaches users and again whenever it changes.

A team picks an AI agent already registered in Obiguard, chooses what to test and starts a run. Red-team runs attack the agent with Moonshot’s attack modules, which rework a seed prompt in ways such as character swaps, look-alike characters and word substitutions to try to get past the agent’s safeguards. Each turn is judged as refused, partial or complied, and the run reports an attack success rate alongside the full transcript. Benchmark runs use Moonshot’s recipes and cookbooks and grade the agent test by test.

The tests run against the agents an organisation already governs in Obiguard, on any model they are connected to, including OpenAI, Anthropic, Azure OpenAI, Amazon Bedrock and Google Gemini, so there is no separate test harness to build or maintain.

Teams can run the same test with Obiguard’s policy enforcement off and on. The raw model and the guarded system are graded the same way, so the difference between the two is evidence that the guardrails work. Every prompt, response and verdict behind a grade is kept, so a reviewer can see why a result came out as it did, and each run records the prompt version and policy set that were tested, so a grade still describes what was tested after the agent changes.

Security and red teams can use it to find weaknesses with repeatable attacks before an adversary does. AI engineers and product teams can check that a new model, prompt or policy change has not made an agent less safe before it ships. Risk, compliance and audit teams get test results with the evidence behind them, alongside the audit trail and ethics review already in Governance AI. Security leaders can show boards and regulators what the guardrails actually change.

Red-teaming and evaluations are part of Governance AI, which is available now. Learn more about AI red-teaming and evaluations or Governance AI.

About Obiguard

Obiguard is an AI risk governance platform founded in Malaysia in 2024 and backed by Antler. It helps organisations inspect AI usage, enforce policy, and keep the audit evidence regulators expect. Learn more at obiguard.ai/about.

Media contact: [email protected]

← All press