For information only. Not advice, and not to be relied on. What that means, in full

AILuminate - MLCommons

https://mlcommons.org/benchmarks/ailuminate/

MLCommons AILuminate — a consortium safety benchmark, and the closest thing to an industry-standard independent safety score
1
Lines added
7
Lines removed
Was read
2026-09-08
2cb7d031d90dcbb3
Changed by
2026-09-10
ccb8b8a2e4cf1e92

1 line added, 7 lines removed

Left column is the version read 2026-09-08, right is 2026-09-10. Unchanged stretches are elided. Both bodies are kept in full and are addressed by the hashes above.
44 unchanged lines
4545 Benchmark Details
4646  
4747 AILuminate general purpose AI chat model benchmark
4848  
49The AILuminate benchmark is designed to evaluate the safety of a fine-tuned LLM general purpose AI chat model. It evaluates the level of potential harm across twelve hazard categories focused on physical hazards, non-physical hazards and contextual hazards.
49+The AILuminate benchmark is designed to evaluate the safety of a fine-tuned LLM general purpose AI chat model. It evaluates the level of potential harm across twelve hazard categories focused on physical hazards, non-physical hazards and contextual hazards.
5050  
5151 General purpose AILuminate AI Chat Model Benchmark supports:
5252  
5353 Single Turn: the current version supports only single turn conversations (a human prompt and a machine response).
30 unchanged lines
8484  
8585 Learn about our AILuminate Safety Benchmark
8686  
8787 Learn about our AILuminate Jailbreak Benchmark
88 
89Want to test your own SUT?
90 
91To inquire about testing your system complete this form.
92 
93Submit a SUT
GovernanceHub — the governance registry and policy router for AI systems