MLCommons — AILuminate benchmark (AI Risk & Reliability working group)
- Type
- Standards body
- Lieu
- International — International (engineering consortium)
- Dernière vérification
- 2026-09-22
- Prochaine vérification
- 2027-03-21
Pas encore traduit — affiché en anglais.
Comment les joindre
- AI Risk & Reliability working group — open to participants
Ce canal est conservé uniquement comme référence. Sa vérification est absente, date de plus de 180 jours, ou la période de contact n'est pas ouverte actuellement ; ne vous y fiez pas avant une nouvelle vérification.
- Get involved
Ce canal est conservé uniquement comme référence. Sa vérification est absente, date de plus de 180 jours, ou la période de contact n'est pas ouverte actuellement ; ne vous y fiez pas avant une nouvelle vérification.
- GitHub — code, taxonomy and prompt-set contributions/issues
- MLCommons Discord
- Public benchmark results
- Homepage
Ce qu'il fait
AILuminate is a standardised AI risk assessment benchmark for chat and vision-language models across 12 hazard categories, built with industry, academia and civil society; at last check 59,624 test prompts, 477 test images and 109 models benchmarked, with expanding agentic, multimodal and jailbreak/security suites. Produces public grades per model.
Évaluation franche
One of the more genuinely participatory technical bodies. If you can demonstrate a hazard category the benchmark misses, the working group is a place where that can actually change an artefact that companies are graded on. Requires technical credibility and sustained attendance.
Préoccupations couvertes
Product-level hazard and reliability: violent crime facilitation, CSAM, hate, self-harm, privacy, IP, defamation, specialised advice. Near-term deployment harms and misuse — NOT loss of control.
Canal gouvernemental
Moderate — AILuminate is cited in policy discussion of evaluation standards and MLCommons engages with NIST and the safety-institute network, but it has no formal mandate.