Baseline
$6,500
5 business days
One agent flow, full 51-scenario suite, single-model run, judge-triangulated score.
AgentBoundary Certified
Adversarial testing against 51 prompt-injection scenarios, execution-based scoring, and a verification package using the open AgentBoundary benchmark.
Certification paths
AI SaaS vendors, security teams deploying internal copilots and agents, and LLM/MLOps platforms that need an independent report on a defined agent workflow.
Use this before enterprise security review, before granting production access to an agent, or when a model upgrade has invalidated prior testing.
The problem
Certification tiers
Three tiers, from a single-flow baseline to a board-ready engagement with bespoke adversarial scenarios.
$6,500
5 business days
One agent flow, full 51-scenario suite, single-model run, judge-triangulated score.
$18,000
10 business days
Up to three agent flows, all four scenario types (RAG, filesystem, browser, code), cross-judge triangulation, one rerun after fixes, procurement memo.
$38,000
15 business days
Up to six flows, full suite plus bespoke attack extensions, executive readout, remediation workshop, quarterly recert option.
Annual recertification costs 55–60% of the initial tier. Targeted retesting after a major model change is scoped to the affected workflows.
What you receive
Overall compliance %, per-attack-family rates, Wilson confidence intervals on every rate, and a 0–100 AgentBoundary Robustness Score.
The failing prompts, execution traces, root-cause patterns, and a severity-ranked remediation backlog written for engineering action.
Methodology, judge-triangulation summary, scope boundaries, explicit limitations, and a signed attestation letter.
12-month verification listing with tier, certification date, verification URL, and revocation terms. Confidential by default; the vendor controls publication.
Annual recertification and targeted retesting after material model or architecture changes. Each result applies to the recorded scope, versions, and test date.
Methodology
The review uses the AgentBoundary benchmark. Reports document the tested workflow, scenario set, scoring method, and limitations so reviewers can understand how each result was reached.
Hard questions, answered
“We already run the open-source benchmark, or our internal red team tests this.”
Independent testing can complement your internal review with a documented scope, recorded results, and an audit trail that customer security teams can examine.
“LLM-as-judge scoring is subjective, or could be wrong.”
We use execution-based scoring with structured rubrics, multiple judges, and Wilson score confidence intervals on measured compliance rates. Model judgments can still be wrong; results must be read with the test scope and limitations.
“A bad score could be used against us.”
All results are confidential. You have sole discretion over whether and when to display the public badge. A private pre-certification window is included; publish only after a remediation pass clears your chosen threshold.
Why it holds
The benchmark is open by design. The certification adds a published methodology, a repeatable adversarial dataset, active-execution scoring, signed validation materials, and a clear audit trail that security and procurement teams can review.
Certification scope
Apply to start with a short scoping call. Turnaround, remediation windows, and any public verification options are confirmed in scope.