AILuminate - MLCommons
https://mlcommons.org/benchmarks/ailuminate/
MLCommons AILuminate — a consortium safety benchmark, and the closest thing to an industry-standard independent safety score
1
Lines added
7
Lines removed
Was read
2026-09-08
2cb7d031d90dcbb3…
Changed by
2026-09-10
ccb8b8a2e4cf1e92…
1 line added, 7 lines removed
Left column is the version read 2026-09-08, right is 2026-09-10. Unchanged stretches are elided. Both bodies are kept in full and are addressed by the hashes above.
⋯ 44 unchanged lines
4545 Benchmark Details
4646
4747 AILuminate general purpose AI chat model benchmark
4848
49−The AILuminate benchmark is designed to evaluate the safety of a fine-tuned LLM general purpose AI chat model. It evaluates the level of potential harm across twelve hazard categories focused on physical hazards, non-physical hazards and contextual hazards.
49+The AILuminate benchmark is designed to evaluate the safety of a fine-tuned LLM general purpose AI chat model. It evaluates the level of potential harm across twelve hazard categories focused on physical hazards, non-physical hazards and contextual hazards.
5050
5151 General purpose AILuminate AI Chat Model Benchmark supports:
5252
5353 Single Turn: the current version supports only single turn conversations (a human prompt and a machine response).
⋯ 30 unchanged lines
8484
8585 Learn about our AILuminate Safety Benchmark
8686
8787 Learn about our AILuminate Jailbreak Benchmark
88−
89−Want to test your own SUT?
90−
91−To inquire about testing your system complete this form.
92−
93−Submit a SUT