v 2.6 · GA5 frameworks · Moonshot red-teaming · nightly ethics scoring

AI risk governance,
built for the way your
organization actually ships AI.

Obiguard Governance AI inspects every prompt, response, and tool-call across your stack — enforces policy, blocks exfiltration, and gives security & legal a single ledger of evidence. It also red-teams your agents against the Project Moonshot catalogue, scores their outputs for bias and fairness every night, and reports coverage against five AI governance frameworks.

LLM providers
40+ supported
Guardrail detectors
11shipped
Frameworks mapped
5self-assessed
Ethics judges
6nightly
01 / The Specimen

A live view from
a customer's production project.

Three violations fired in the past minute. One blocked by a Policy Set. Two routed to the Review Queue. All written to the Audit Log — searchable, exportable, and evidence-ready.
streaming · 12,402 ev/hr

Violations

last 60 minutes · 14 violations · 3 blocked · 5 flagged for review
Live
History
Export
Requests
12,402
+4.1%
Threats blocked
38
High: 6
Avg latency
18ms
p50
Active policy sets
47
3 drafts
Time
Event
Policy
Severity
Action
14:02:09
SSN in completion → support-agent · gpt-4.1
PII / US-SSN
HIGH
REDACT
14:01:58
Tool out-of-scope · db.query → payroll
Tool / Scope
HIGH
BLOCK
14:01:54
Prompt injection · act-as-developer-mode
Injection
HIGH
BLOCK
14:01:42
Off-allow-list model · llama-3.1-405b
Model / AL
MED
REVIEW
14:00:48
Customer record leak → rag-search
DLP / Customer
MED
REDACT
13:59:50
Toxicity 0.81 in agent reply
Content / Tox
LOW
LOGGED
02 / Capabilities

Nine capabilities. One platform.
Every AI agent your org runs.

Obiguard wraps every AI agent your organisation operates — monitoring usage, enforcing policies, routing violations for human review, red-teaming agents before they ship, and maintaining the framework-mapped evidence your compliance team needs.
01 — MONITOR

Dashboard & violations

A live project dashboard shows token usage, cost, request volume, errors, threats blocked, and average latency. The Violations view surfaces every policy breach — filtered, searchable, and exportable.

Real-time · threats blocked · p50 latency
02 — REVIEW

Human-in-the-loop review queue

Violations flagged for human judgment land in the Review Queue. Assign, annotate, and resolve — with a full audit trail from the original event to the decision.

Review Queue · assign & resolve
03 — ENFORCE

Policy Sets

Group enforcement rules into Policy Sets and assign them to AI Agents. Each set defines criteria, action (block / flag / allow), and whether violations route to the Review Queue.

Policy Sets · per-agent enforcement
04 — REGISTER

AI Use Cases & Agents

Maintain a living registry of every AI Use Case and Agent in your organisation. Link each agent to its policy set — so governance follows the workload, not the calendar.

Use Cases · agents & prompts
05 — LOG

Audit Log

Every prompt, response, tool-call, and policy decision is written to an Audit Log. Export it to CSV for your GRC stack, or hand the auditor a timestamped record.

Audit Log · CSV export
06 — GOVERN

Controls & Criteria

Define organisation-wide Controls and the Criteria that trigger them. Controls carry an approval workflow with segregation of duties — a submitter cannot approve their own control.

Controls · approval workflow
07 — MEASURE

Ethics & Bias report

A nightly job samples the previous day’s traces and scores each one against six LLM judge criteria — toxicity, demographic bias, sentiment consistency, fairness, misinformation, and privacy leakage. Anything below 60 is flagged for review.

6 judges · nightly · never blocks traffic
08 — TEST

Evaluations & red-teaming

Benchmark and red-team your registered agents against the Project Moonshot catalogue from AI Verify Foundation — safety, bias, jailbreak resistance, privacy leakage, and capability. Every run is graded and retained as governance evidence.

Project Moonshot · cookbooks & attack modules
09 — MAP

Compliance frameworks

See how your policies, controls, criteria, use cases, and violations line up against NIST AI RMF, ISO/IEC 42001, OWASP LLM Top 10, AISCF, and Singapore’s MGF for GenAI. Each category shows covered / partial / not covered, with the evidence behind it.

5 frameworks · CSV export

Go deeper on any of them: inspection, Policy Sets, the Audit Ledger, Ethics & Bias reporting, Project Moonshot evaluations, and framework mapping.

For Security

Make every AI agent a governed workload.

  • Violations dashboard with real-time threats blocked
  • Policy Sets block or flag unsafe agent behaviour
  • Red-team agents with Moonshot attack modules before they ship
  • Audit Log exports to CSV for your SIEM or GRC pipeline
→ For CISOs & security platform teams
For Risk & Legal

Continuous evidence, not quarterly screenshots.

  • Coverage against NIST AI RMF, ISO 42001, OWASP LLM Top 10, AISCF, MGF for GenAI
  • Nightly Ethics & Bias scoring as fairness evidence
  • Audit Log — every decision timestamped
  • Export as evidence for internal AUP audits and self-assessment
→ For Risk, Legal & Compliance
For AI Teams

Register, govern, and iterate on your AI agents.

  • AI Use Case & Agent registry — one source of truth
  • Policy Sets assign enforcement rules per agent
  • Benchmark a prompt change against the Moonshot catalogue before rollout
  • Review Queue routes edge cases to human judgment
→ For AI product & platform teams
03 / How it works

Live in 14 days.
Assurance from day one.

No model retraining. No migration. Obiguard slots in front of your existing AI gateway or speaks AWS/Azure/GCP-native APIs.
STEP 01 · CONNECT

30 minutes

Create a project, generate an Access Key, and route your AI agents through Obiguard. Works with OpenAI, Anthropic, Bedrock, Vertex, Azure OpenAI, and any OpenAI-compatible endpoint.

# route via Obiguard
$ export OPENAI_BASE_URL=\
  https://gw.obiguard.com/v1
→ project connected
STEP 02 · REGISTER

Day one

Add your AI Use Cases and register each AI Agent in the platform. Link agents to the business function they serve — so every governance decision has full context.

AI Use Cases registered
3 agents linked
1 project · t77
→ registry complete
STEP 03 · ENFORCE

When you're ready

Build a Policy Set — define your Criteria, set the action (block or flag), and assign the set to your agents. Violations surface immediately in the dashboard.

policy_set: no-pii-in-responses
action: block
agent: support-agent
review_queue: true
STEP 04 · ASSURE

Continuously

Red-team each agent against the Moonshot catalogue, let the nightly Ethics & Bias job score live traffic, and read framework coverage off one page. Violations still land in the Audit Log and Review Queue — export the whole lot as evidence.

eval: singapore-pofma-statements
grade → E · 2 failing of 2
ethics → 95 · good
nist-ai-rmf — 40% covered
04 / What ships today

The platform, not the pitch deck.

Everything below is running in the product today — including Project Moonshot evaluations, the nightly Ethics & Bias report, and framework coverage across five standards. Inspection runs at the gateway layer: no retraining, no access to your weights or training pipeline.

11
Guardrail detectors

Jailbreak/injection, PII, NSFW, toxicity, gibberish, profanity, ban list, competitor mentions, and format guards — plus keyword, regex, and LLM-judge criteria you write yourself.

40+
LLM providers supported

OpenAI, Anthropic, Bedrock, Azure OpenAI, Vertex AI, and more — one gateway in front of all of them.

5
Frameworks mapped

NIST AI RMF, ISO/IEC 42001, OWASP LLM Top 10, Malaysia’s AISCF, and Singapore’s MGF for GenAI — as a self-assessment aid, not a certification.

6
Ethics judges, nightly

Toxicity, demographic bias, sentiment consistency, fairness, misinformation, and privacy leakage — scored on a random sample of yesterday’s traces.

05 / Frameworks

Two global. One security.
Two for Southeast Asia.

Coverage is estimated from your own platform activity and shown per category as covered, partial, or not covered — with the evidence behind each call. It is a self-assessment aid for audit preparation, not a certification.
NIST AI RMF

NIST AI Risk Management Framework

The US reference vocabulary for AI risk — Govern, Map, Measure, Manage. What most enterprise risk teams have already adopted internally.

US National Institute of Standards and Technology · 4 functions · 10 categories mapped
ISO/IEC 42001

ISO/IEC 42001 — AI Management System

The only certifiable standard here — an AI management system in the same family as ISO 27001, audited by an accredited body.

ISO / IEC · 5 clauses mapped (6 – 10)
OWASP LLM Top 10

OWASP Top 10 for LLM Applications

How LLM applications actually get broken — injection, disclosure, excessive agency. The list your security team already works from.

OWASP Foundation · 10 risk categories
AISCF

AI Systems Cyber Security Framework

Malaysia’s national framework from NACSA — 42 security activities across a seven-phase AI system lifecycle, from Inception to Retirement.

NACSA — National Cyber Security Agency of Malaysia · 7 lifecycle phases · 42 security activities
MGF for GenAI

Model AI Governance Framework for Generative AI

Singapore’s baseline for trusted generative AI from IMDA and AI Verify Foundation — nine dimensions, five of them a deployer can evidence.

IMDA & AI Verify Foundation, Singapore · 9 dimensions · 5 platform-evidenceable
YOUR AUP

Your own Acceptable Use Policy

Define your own Controls and Criteria with custom labels and evidence mapping — the same coverage view, scored against the policy your organisation actually wrote.

Custom · controls & criteria
What each framework asks for, and how Obiguard evidences it →
06 / FAQ

Common
questions.

Don't see what you're looking for? Our solutions engineers respond within one business day.

Talk to an SE →
Does Obiguard see our customers' data?[01]
In SaaS mode, content is processed in-memory to make a policy decision — only the decision and metadata are persisted to the Audit Log. Private and on-prem deployment options are available for enterprise customers; talk to a solutions engineer about your requirements.
Does this add latency to model calls?[02]
Inspection runs on the request path, so it adds some latency — how much depends on which detectors are enabled for your policy. We're happy to benchmark your specific configuration during a pilot.
How is Obiguard different from a model gateway?[03]
Gateways route traffic. Obiguard governs it. We sit in front of your gateway or SDK and add the policy enforcement, violation tracking, review workflow, and audit trail that routing tools leave out of scope.
Which compliance frameworks do you support?[04]
Five today. NIST AI RMF 1.0 and ISO/IEC 42001:2023 are what a multinational customer or auditor will ask for — the second is the only certifiable standard in the set. The OWASP Top 10 for LLM Applications (2025) is the security list your penetration testers already work from. AISCF is Malaysia's AI Systems Cyber Security Framework, published by NACSA: 42 security activities organised across a seven-phase system lifecycle, from Inception through Operation & Monitoring to Retirement.MGF for GenAI is Singapore's Model AI Governance Framework for Generative AI, from IMDA and the AI Verify Foundation — nine dimensions, of which four are ecosystem commitments aimed at model developers and policymakers rather than deployers, so that score is capped at 56% by design. Coverage is estimated from your own platform activity, category by category, as a self-assessment aid for audit preparation — not a certification. See what each framework asks for →
What is Project Moonshot, and how do evaluations work?[05]
Project Moonshot is AI Verify Foundation’s open testing toolkit for LLMs. Obiguard runs its catalogue of cookbooks and recipes — safety, bias, jailbreak resistance, privacy leakage, and capability — plus its red-teaming attack modules, against the AI Agents you have registered. Pick a target and a system prompt, run the benchmark, and every recipe comes back graded and retained as governance evidence. Grading scales differ per recipe (some invert, some band into risk levels), so we always render the grade alongside its scale rather than a bare score.
How does the Ethics & Bias report work?[06]
A job runs at 03:00 UTC each day and randomly samples roughly 17% of the previous day’s traces. Each sampled trace is scored 0–100 against six LLM judge criteria — toxicity, demographic bias, sentiment consistency, fairness, misinformation, and privacy leakage. Anything scoring under 60 on any evaluator is flagged for review. It runs asynchronously and never blocks live traffic, and the report shows a 30-day trend, per-agent scores, and drill-down into every flagged trace.
How do Policy Sets and Criteria work together?[07]
Criteria define what triggers a policy event — a detector match, a keyword, a model call outside scope. Policy Sets group Criteria and assign them to AI Agents, with an action (block or flag) and an optional route to the Review Queue.
What is the Review Queue?[08]
Violations that need human judgment — edge cases, high-risk content, new agent behaviour — are routed to the Review Queue. Reviewers can annotate, approve, or escalate, and every decision is written to the Audit Log.
07 / Get started

Bring the same rigor you apply
to data and identity — to the AI
in your stack.

A 20-minute call with a solutions engineer is enough to scope your pilot. Most customers are enforcing policy in production within two weeks.

Read the architectureDownload brochure (PDF) ↓