Platform · EvaluationsPowered by Project Moonshot — AI Verify Foundation

Red-team your AI agents
before an attacker
does it for you.

Benchmark and red-team every registered AI Agent against the Project Moonshot test catalogue — safety, bias, jailbreak resistance, privacy leakage, and capability. Results are graded, retained as governance evidence, and mapped straight onto the framework categories that ask for adversarial testing.

01 / How evaluations work

Pick a target. Pick a test. Get a grade.

Targets are the AI Agents you have already registered, so Obiguard resolves the provider connector, model, and system prompt for you — the portal never handles a provider credential.
01 — CATALOGUE

Project Moonshot cookbooks & recipes

The open testing catalogue from AI Verify Foundation, served live and pinned to a known upstream commit. Browse cookbooks by category, read what each recipe measures, and see how many prompts a run will send before you start it.

Cookbooks · recipes · attack modules
02 — TARGET

Test the agent you actually ship

Run against a registered AI Agent using its current system prompt, another prompt from the registry, or a custom one — so you can benchmark a prompt change before it reaches production.

Agent · registry prompt · custom prompt
03 — RED TEAM

Adversarial attack modules

Beyond fixed benchmarks, Moonshot attack modules probe the agent adaptively — each module owns its own termination logic rather than a fixed prompt count.

Adversarial · session-based
04 — GRADE

Grades, read on the right scale

Moonshot recipes do not share one scale — most are higher-is-better, the toxicity and AdvGLUE recipes invert it, and the MLCommons recipes band into risk levels instead of letters. Obiguard always renders the grade with its scale, never a bare score.

A–E · inverted · risk bands
05 — TRIAGE

Failing recipes, surfaced

The Evaluations page reads posture from the latest completed run per target, so a stale run against a replaced model cannot drag the headline number around. Failing recipes are listed as the findings to action.

Latest run · per target
06 — EVIDENCE

Retained as governance evidence

Every run — target, prompt under test, catalogue version, grade, and failing recipes — is retained against the project and counts toward the framework categories covering adversarial testing and trustworthy-characteristic evaluation.

NIST MEASURE 2.7 · MGF Testing & Assurance
AIVF
Project Moonshot

AI Verify Foundation's open LLM testing toolkit, run inside your Obiguard project.

5
Test dimensions

Safety, bias, jailbreak resistance, privacy leakage, and capability.

A–E
Graded, with scale

Every recipe reported against its own grading scale — inverted and risk-band recipes included.

0
Credentials handled

Targets resolve through registered agents; provider tokens are injected server-side per run.

03 / FAQ

Evaluations,
explained.

What Project Moonshot is, how a benchmark differs from a red-team run, and how to read a grade.

Talk to an SE →

What is Project Moonshot?[01]

Project Moonshot is the open LLM testing toolkit from the AI Verify Foundation, Singapore’s AI governance testing body. It provides a catalogue of benchmarks — organised as cookbooks and recipes covering safety, bias, jailbreak resistance, privacy leakage, and capability — plus adversarial red-teaming attack modules. Obiguard runs that catalogue inside your own project, against the AI Agents you have registered.

What is the difference between a benchmark run and red-teaming?[02]

A benchmark run sends a fixed catalogue of prompts — a cookbook or a single recipe — and grades the answers, so results are comparable between runs and between models. Red-teaming uses attack modules that probe the agent adaptively; each module owns its own termination logic rather than a fixed prompt count, so it explores rather than measures. Most teams use benchmarks to track regression over time and attack modules to find something new.

What do the A to E grades mean?[03]

Grades come from each recipe’s own grading scale, and those scales are not all the same. Most recipes are higher-is-better, where A is 80–100 and E is 0–19. The AdvGLUE and toxicity recipes invert that, because their raw score measures attack success or toxicity rather than performance. The MLCommons recipes do not use letters at all and band into Low Risk through High Risk instead. Obiguard always renders the grade alongside its scale, because a raw score without its scale is meaningless.

What can I run an evaluation against?[04]

Any AI Agent registered in the project. Obiguard resolves the agent to its provider connector and model and injects the credential server-side per run, so the portal never handles a provider token. You choose whether to test the agent’s current system prompt, another prompt from the registry, or a custom one — which is how teams benchmark a prompt change before it reaches production.

Do evaluation runs affect production traffic?[05]

No. Runs are dispatched as separate benchmarking or red-teaming sessions against the target model; they do not sit on your live request path and do not consume your policy enforcement budget. The Evaluations page reads posture from the latest completed run per target, so an old run against a since-replaced model cannot drag the headline grade around.

Do evaluation results count as compliance evidence?[06]

Yes. Each run is retained against the project with its target, the prompt under test, the catalogue version, the resulting grade, and any failing recipes. That record is what framework categories covering adversarial testing ask for — NIST AI RMF MEASURE 2.7, AISCF CR06 security assessment and testing, and the testing side of ISO/IEC 42001 Clause 9.
04 / Related

Testing is one half of assurance.

Evaluations tell you how an agent behaves under test. The nightly Ethics & Bias report tells you how it behaved on real traffic, and framework mapping turns both into coverage against NIST AI RMF, ISO/IEC 42001, AISCF and MGF for GenAI. Enforcement happens upstream in inspection and Policy Sets — all of it on the Governance AI page.