MLCommons — AILuminate benchmark (AI Risk & Reliability working group)
- Type
- Standards body
- Place
- International — International (engineering consortium)
- Last checked
- 2026-09-22
- Next check due
- 2027-03-21
Ways to reach them
- AI Risk & Reliability working group — open to participants
This route is retained as reference only. Its verification is missing, more than 180 days old, or the contact window is not currently open; do not rely on it until it is re-checked.
- Get involved
This route is retained as reference only. Its verification is missing, more than 180 days old, or the contact window is not currently open; do not rely on it until it is re-checked.
- GitHub — code, taxonomy and prompt-set contributions/issues
- MLCommons Discord
- Public benchmark results
- Homepage
What it does
AILuminate is a standardised AI risk assessment benchmark for chat and vision-language models across 12 hazard categories, built with industry, academia and civil society; at last check 59,624 test prompts, 477 test images and 109 models benchmarked, with expanding agentic, multimodal and jailbreak/security suites. Produces public grades per model.
Honest assessment
One of the more genuinely participatory technical bodies. If you can demonstrate a hazard category the benchmark misses, the working group is a place where that can actually change an artefact that companies are graded on. Requires technical credibility and sustained attendance.
Concerns it covers
Product-level hazard and reliability: violent crime facilitation, CSAM, hate, self-harm, privacy, IP, defamation, specialised advice. Near-term deployment harms and misuse — NOT loss of control.
Government channel
Moderate — AILuminate is cited in policy discussion of evaluation standards and MLCommons engages with NIST and the safety-institute network, but it has no formal mandate.