Benchmark and red-team every registered AI Agent against the Project Moonshot test catalogue — safety, bias, jailbreak resistance, privacy leakage, and capability. Results are graded, retained as governance evidence, and mapped straight onto the framework categories that ask for adversarial testing.
The open testing catalogue from AI Verify Foundation, served live and pinned to a known upstream commit. Browse cookbooks by category, read what each recipe measures, and see how many prompts a run will send before you start it.
Run against a registered AI Agent using its current system prompt, another prompt from the registry, or a custom one — so you can benchmark a prompt change before it reaches production.
Beyond fixed benchmarks, Moonshot attack modules probe the agent adaptively — each module owns its own termination logic rather than a fixed prompt count.
Moonshot recipes do not share one scale — most are higher-is-better, the toxicity and AdvGLUE recipes invert it, and the MLCommons recipes band into risk levels instead of letters. Obiguard always renders the grade with its scale, never a bare score.
The Evaluations page reads posture from the latest completed run per target, so a stale run against a replaced model cannot drag the headline number around. Failing recipes are listed as the findings to action.
Every run — target, prompt under test, catalogue version, grade, and failing recipes — is retained against the project and counts toward the framework categories covering adversarial testing and trustworthy-characteristic evaluation.
AI Verify Foundation's open LLM testing toolkit, run inside your Obiguard project.
Safety, bias, jailbreak resistance, privacy leakage, and capability.
Every recipe reported against its own grading scale — inverted and risk-band recipes included.
Targets resolve through registered agents; provider tokens are injected server-side per run.
What Project Moonshot is, how a benchmark differs from a red-team run, and how to read a grade.
Talk to an SE →Evaluations tell you how an agent behaves under test. The nightly Ethics & Bias report tells you how it behaved on real traffic, and framework mapping turns both into coverage against NIST AI RMF, ISO/IEC 42001, AISCF and MGF for GenAI. Enforcement happens upstream in inspection and Policy Sets — all of it on the Governance AI page.