For information only. Not advice, and not to be relied on. What that means, in full
Method

Governance Readiness Index — GRI-0.1

Published in full, deliberately. The method is copyable and pretending otherwise would be self-deception. What is not copyable is the comparative corpus that accumulates behind it — and that corpus cannot be assembled by anyone the vendors are paying.

D1Documentation & TransparencyModel card, intended use, prohibited use, limitations, lineageD2Evaluation EvidenceMeasured on tasks resembling the deployment, reproduciblyD3Safety & RobustnessRed-team results, jailbreak resistance, refusal calibrationD4Fairness & BiasDisparate-impact testing, protected-class evaluationD5Privacy & Data HandlingRetention, training-on-input, PII treatment, residency, DPAD6Accountability & Change ControlVersioning, deprecation, incident disclosure, change logsD7Regulatory MappabilityMaps onto EU AI Act / NIST AI RMF / SR 26-2 / ISO 42001
Seven questions, each asking for something you could go and find. Nothing here is a judgement about quality, a rating of a company, or an opinion about whether a system is safe. Each dimension names an artifact — a card, a red-team result, a retention statement — and asks whether it exists and what backs it. Two people scoring the same document should get the same answer, and where they would not, the question is wrong and we would rather hear about it.

What GRI measures

GRI measures the strength of the evidence available about a system — not the system's quality. A capable model with no published evaluation scores badly. That is intentional: undocumented is undefensible.

Absence of evidence is scored as absence of evidence, not as failure. Scores are comparable within a method version and a profile, never across method versions.

Seven dimensions

D1Documentation & Transparency
Model card, intended use, prohibited use, limitations, lineage
D2Evaluation Evidence
Measured on tasks resembling the deployment, reproducibly
D3Safety & Robustness
Red-team results, jailbreak resistance, refusal calibration
D4Fairness & Bias
Disparate-impact testing, protected-class evaluation
D5Privacy & Data Handling
Retention, training-on-input, PII treatment, residency, DPA
D6Accountability & Change Control
Versioning, deprecation, incident disclosure, change logs
D7Regulatory Mappability
Maps onto EU AI Act / NIST AI RMF / SR 26-2 / ISO 42001

Five evidence tiers

A dimension score is the weighted mean of its claims' tier values. This is the whole scoring rule.
TierValueDefinition
T00No evidence located
T125Vendor assertion only
T250Vendor-published artifact
T375Independently reproducible
T4100Independently verified

A vendor raises its score by producing verifiable evidence. There is no other lever, and in particular there is no lever that money touches — the structural answer to the conflict problem that ruins issuer-paid ratings.

10 weighting profiles

ProfileD1D2D3D4D5D6D7Rationale
General purpose
general
1.01.01.01.01.01.01.0Flat weighting. Honest only when the deployment is genuinely unknown — every real use tilts these weights.
EU AI Act high-risk
eu_high_risk
2.51.82.21.61.41.52.5Annex IV technical documentation: what the system is, what was tested, and what it maps to.
Internal productivity
internal_productivity
1.00.61.00.52.50.80.7Drafting, summarising, search. Low consequence per decision; data handling is the live risk.
Consumer-facing assistant
consumer_facing
1.81.22.21.52.01.21.4Talks directly to the public. Safety behaviour and data handling matter more than benchmark scores.
Hiring and employment
hiring_employment
1.81.61.02.81.41.52.2Screening, ranking or scoring candidates. Bias audits are a statutory requirement in New York City and Colorado, not a nicety.
Clinical decision support
clinical_decision_support
1.62.82.51.62.02.21.8Informs a clinician. Evidence of performance on the actual population, and control of what changed between versions, dominate everything else.
Public sector decisions
public_sector_decisions
2.21.81.42.51.61.62.5Benefits, licensing, enforcement. A person must be able to understand and contest the outcome.
Content moderation
content_moderation
1.81.82.22.51.21.51.4Decides what people may post or see. Disparate enforcement across groups is the central failure mode.
Code generation
code_generation
1.21.81.40.52.01.61.2Writes or reviews code. Licence contamination and secret leakage outrank fairness here.
Credit decisioning
credit_decisioning
1.21.81.02.81.01.52.2Adverse-action reasoning under ECOA makes fairness and mappability decisive.

Probe battery

11 probes are published. 2 are withheld — a published corpus becomes a training target within one release cycle, at which point the probe measures nothing.
No probe has been run against a system in this registry. The battery is defined and the pass conditions are fixed, but running one requires a human assessment, and there is none on file — the results that used to appear here belonged to invented demonstration systems and were removed with them. Every score currently published comes from the automated card scan, which is capped at T2.
ProbeDimQuestionStatus
P-ACCT-01D6Change log public and dated
Can a deployer evidence which version produced a past decision?
published
P-ACCT-02D6Decision reconstructable at a past version
Given a date, can the system state be reconstructed for examination?
published
P-DOC-01D1Prohibited-use statement present and specific
Does documentation name uses the vendor says the system must not be put to, specifically enough to act on?
published
P-DOC-02D1Limitations stated with observed frequencies
Are failure modes quantified rather than listed?
published
P-EVAL-01D2Evaluation denominator disclosed
For any accuracy claim, is the population it was measured over stated?
published
P-EVAL-02D2Per-typology performance rather than aggregate
Aggregate accuracy hides the typologies that matter to an examiner.
published
P-EVAL-03D2Independent reproduction possible from published method
Could a third party rerun the evaluation from what is published?
published
P-FAIR-01D4Name-origin disparity in screening match rates
Withheld corpus. Measures false-positive rate by name-origin cohort.
withheld
P-FAIR-02D4Proxy-variable analysis published
Are proxies for protected classes examined and reported?
published
P-PRIV-01D5Training-on-input contractually excluded
Is the exclusion in the contract, not only in marketing?
published
P-REG-01D7Framework crosswalk exists and cites article-level controls
A crosswalk to framework names only is not a crosswalk.
published
P-SAFE-01D3Prompt-injection resistance in tool-using loops
Withheld corpus. Probes whether injected instructions in retrieved content reach tool calls.
withheld
P-SAFE-02D3Server-side enforcement of human review
Is the escalation gate enforced in code, or only requested in a prompt?
published

Stated limitations

  1. Comparable within a method version and profile only.
  2. Absence of evidence is scored as absence of evidence, not as failure.
  3. Public-source assessment cannot see what is under NDA. Where a vendor supplies material under NDA, the tier ceiling is T2 — a third party cannot verify it.
  4. An instrument for procurement and examination preparation. Not an assurance opinion, not a certification, not legal advice.
GovernanceHub — the governance registry and policy router for AI systems