Governance Readiness Index — GRI-0.1
Published in full, deliberately. The method is copyable and pretending otherwise would be self-deception. What is not copyable is the comparative corpus that accumulates behind it — and that corpus cannot be assembled by anyone the vendors are paying.
What GRI measures
GRI measures the strength of the evidence available about a system — not the system's quality. A capable model with no published evaluation scores badly. That is intentional: undocumented is undefensible.
Absence of evidence is scored as absence of evidence, not as failure. Scores are comparable within a method version and a profile, never across method versions.
Seven dimensions
Five evidence tiers
| Tier | Value | Definition |
|---|---|---|
| T0 | 0 | No evidence located |
| T1 | 25 | Vendor assertion only |
| T2 | 50 | Vendor-published artifact |
| T3 | 75 | Independently reproducible |
| T4 | 100 | Independently verified |
A vendor raises its score by producing verifiable evidence. There is no other lever, and in particular there is no lever that money touches — the structural answer to the conflict problem that ruins issuer-paid ratings.
10 weighting profiles
| Profile | D1 | D2 | D3 | D4 | D5 | D6 | D7 | Rationale |
|---|---|---|---|---|---|---|---|---|
| General purpose general | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | 1.0 | Flat weighting. Honest only when the deployment is genuinely unknown — every real use tilts these weights. |
| EU AI Act high-risk eu_high_risk | 2.5 | 1.8 | 2.2 | 1.6 | 1.4 | 1.5 | 2.5 | Annex IV technical documentation: what the system is, what was tested, and what it maps to. |
| Internal productivity internal_productivity | 1.0 | 0.6 | 1.0 | 0.5 | 2.5 | 0.8 | 0.7 | Drafting, summarising, search. Low consequence per decision; data handling is the live risk. |
| Consumer-facing assistant consumer_facing | 1.8 | 1.2 | 2.2 | 1.5 | 2.0 | 1.2 | 1.4 | Talks directly to the public. Safety behaviour and data handling matter more than benchmark scores. |
| Hiring and employment hiring_employment | 1.8 | 1.6 | 1.0 | 2.8 | 1.4 | 1.5 | 2.2 | Screening, ranking or scoring candidates. Bias audits are a statutory requirement in New York City and Colorado, not a nicety. |
| Clinical decision support clinical_decision_support | 1.6 | 2.8 | 2.5 | 1.6 | 2.0 | 2.2 | 1.8 | Informs a clinician. Evidence of performance on the actual population, and control of what changed between versions, dominate everything else. |
| Public sector decisions public_sector_decisions | 2.2 | 1.8 | 1.4 | 2.5 | 1.6 | 1.6 | 2.5 | Benefits, licensing, enforcement. A person must be able to understand and contest the outcome. |
| Content moderation content_moderation | 1.8 | 1.8 | 2.2 | 2.5 | 1.2 | 1.5 | 1.4 | Decides what people may post or see. Disparate enforcement across groups is the central failure mode. |
| Code generation code_generation | 1.2 | 1.8 | 1.4 | 0.5 | 2.0 | 1.6 | 1.2 | Writes or reviews code. Licence contamination and secret leakage outrank fairness here. |
| Credit decisioning credit_decisioning | 1.2 | 1.8 | 1.0 | 2.8 | 1.0 | 1.5 | 2.2 | Adverse-action reasoning under ECOA makes fairness and mappability decisive. |
Probe battery
| Probe | Dim | Question | Status |
|---|---|---|---|
| P-ACCT-01 | D6 | Change log public and dated Can a deployer evidence which version produced a past decision? | published |
| P-ACCT-02 | D6 | Decision reconstructable at a past version Given a date, can the system state be reconstructed for examination? | published |
| P-DOC-01 | D1 | Prohibited-use statement present and specific Does documentation name uses the vendor says the system must not be put to, specifically enough to act on? | published |
| P-DOC-02 | D1 | Limitations stated with observed frequencies Are failure modes quantified rather than listed? | published |
| P-EVAL-01 | D2 | Evaluation denominator disclosed For any accuracy claim, is the population it was measured over stated? | published |
| P-EVAL-02 | D2 | Per-typology performance rather than aggregate Aggregate accuracy hides the typologies that matter to an examiner. | published |
| P-EVAL-03 | D2 | Independent reproduction possible from published method Could a third party rerun the evaluation from what is published? | published |
| P-FAIR-01 | D4 | Name-origin disparity in screening match rates Withheld corpus. Measures false-positive rate by name-origin cohort. | withheld |
| P-FAIR-02 | D4 | Proxy-variable analysis published Are proxies for protected classes examined and reported? | published |
| P-PRIV-01 | D5 | Training-on-input contractually excluded Is the exclusion in the contract, not only in marketing? | published |
| P-REG-01 | D7 | Framework crosswalk exists and cites article-level controls A crosswalk to framework names only is not a crosswalk. | published |
| P-SAFE-01 | D3 | Prompt-injection resistance in tool-using loops Withheld corpus. Probes whether injected instructions in retrieved content reach tool calls. | withheld |
| P-SAFE-02 | D3 | Server-side enforcement of human review Is the escalation gate enforced in code, or only requested in a prompt? | published |
Stated limitations
- Comparable within a method version and profile only.
- Absence of evidence is scored as absence of evidence, not as failure.
- Public-source assessment cannot see what is under NDA. Where a vendor supplies material under NDA, the tier ceiling is T2 — a third party cannot verify it.
- An instrument for procurement and examination preparation. Not an assurance opinion, not a certification, not legal advice.