How the rating works
The AI Agent Governance Assessment is deliberately modeled on a credit rating: a consistent, published methodology, applied the same way to every organization, producing a grade that means the same thing everywhere. This page is that methodology.
Why publish it
A score issued by a company that benefits when you grade low has a credibility problem — unless the methodology is public, the criteria are concrete, and some organizations genuinely score well. All three are true here. The rating is worth taking precisely because it is real: it is evidence-anchored, benchmarked, and some programs clear the investment-grade floor on the first pass.
Three pillars
The rating scores an AI agent program on the three things most organizations cannot answer with confidence.
Governance — 40%
Can you see, secure, and control what your agents do? Oversight, identity and least privilege, containment, plus the security of the humans and infrastructure around the system.
Heaviest weight — this is where existential risk livesOptimization — 30%
Can you trust that your agents are right? Grounding, evaluation, hallucination control, uncertainty handling, and containment of cascading errors between agents.
Accuracy without control is exposureMAS Maturity — 30%
Is the program run deliberately? Executive ownership, a complete agent inventory, lifecycle management, and literacy in multi-agent failure modes.
You cannot govern what you cannot seeAn evidence-anchored scale
Every item is scored on the same five-level maturity scale, mirroring recognized AI-governance maturity models. In a verified rating, claims require artifacts — a claim with nothing to show for it is capped at Level 1 no matter how it is described. That single rule is what separates a rating from a survey.
| Level | Name | What it means |
|---|---|---|
| 0 | Absent | No capability. Nothing exists. |
| 1 | Ad hoc | Informal, inconsistent, undocumented — lives in people's heads or prompts. |
| 2 | Defined | Documented and partially implemented. |
| 3 | Managed | Implemented consistently across most agents; monitored. |
| 4 | Optimized | Implemented, measured, automated, and continuously improved. |
Items are also risk-weighted: catastrophic-risk controls count three times as much as good-hygiene items. A missing kill switch moves your rating far more than an immature feedback loop — by design.
A sample item
“When an agent misbehaves, can you isolate it, stop it, and roll back what it did — and does every agent have a named human owner?”
Level 0: no kill switch, no named owners. · Level 2: agents can be stopped manually, rollback unclear, partial ownership. · Level 4: tested containment and rollback; every agent has a named owner and an incident path.
In a verified rating, this item is scored against artifacts: a documented containment procedure, a rollback runbook, and the agent-owner registry.
The grade ladder
| Rating | Grade | Tier | Read |
|---|---|---|---|
| 90–100 | AAA | Fortified | Governed, measured, improving. Reference-grade. |
| 80–89 | AA | Resilient | Strong controls; gaps are refinements. |
| 70–79 | A | Established | Solid foundation, consistently applied. |
| 60–69 | BBB | Developing | Foundations in place but uneven. Investment-grade floor. |
| 50–59 | BB | Vulnerable | Real gaps; controls don't yet cover the exposure. |
| 35–49 | B | Exposed | Minimal controls; significant live risk. |
| 0–34 | CCC | Critical | Little to no governance. Urgent. |
BBB is the investment-grade floor — the line, borrowed from credit ratings, between “governed enough to scale” and speculative.
Exposure bends the grade
Five intake dimensions — deployment stage, footprint, process criticality, autonomy, and blast radius — measure how much is at stake right now. Exposure isn't good or bad, but it changes what a score means: high live exposure with weak governance cannot be investment grade, so the model caps the final grade in those cases — the same logic a rating agency applies to an over-leveraged borrower. The same low rating reads as unprepared at low exposure and exposed at high exposure. Opposite stories, opposite urgency.
Standards basis
Every item maps to one or more recognized frameworks: the NIST AI Risk Management Framework, ISO/IEC 42001 (the AI management-system standard), the OWASP Top 10 for LLM and Agentic Applications, and MITRE ATLAS for adversarial threats. The rating is an application of accepted standards — not an opinion.
Preliminary vs. verified
The free online assessment is self-reported: it applies this scoring architecture to your own answers and returns a preliminary rating instantly. A verified rating applies the evidence rule — our team reviews actual artifacts against the full rubric and delivers a benchmarked scorecard with a remediation summary, typically within 24–72 hours. Verification is free.