Français

Organismes ayant un pouvoir

MLCommons — AILuminate benchmark (AI Risk & Reliability working group)

Type
Standards body
Lieu
International — International (engineering consortium)
Dernière vérification
2026-09-22
Prochaine vérification
2027-03-21

Pas encore traduit — affiché en anglais.

Comment les joindre

  • AI Risk & Reliability working group — open to participants · Official membership list
    Vérification expirée · 2026-09-22

    Ce canal est conservé uniquement comme référence. Sa vérification est absente, date de plus de 180 jours, ou la période de contact n'est pas ouverte actuellement ; ne vous y fiez pas avant une nouvelle vérification.

    Researchers and practitioners; working-group participation is the substantive route

    MLCommons working groups run open calls and mailing lists; this is a genuine build-it-with-us channel.

  • Get involved · Official membership list
    Vérification expirée · 2026-09-22

    Ce canal est conservé uniquement comme référence. Sa vérification est absente, date de plus de 180 jours, ou la période de contact n'est pas ouverte actuellement ; ne vous y fiez pas avant une nouvelle vérification.

    Anyone

  • GitHub — code, taxonomy and prompt-set contributions/issues · Written submission
    Vérifié · 2026-09-22

    Anyone

  • MLCommons Discord · Web form
    Vérifié · 2026-09-22

    Anyone

  • Public benchmark results · Web form
    Vérifié · 2026-09-22

    Read-only

  • Homepage · Homepage
    Non contrôlé

Ce qu'il fait

AILuminate is a standardised AI risk assessment benchmark for chat and vision-language models across 12 hazard categories, built with industry, academia and civil society; at last check 59,624 test prompts, 477 test images and 109 models benchmarked, with expanding agentic, multimodal and jailbreak/security suites. Produces public grades per model.

Évaluation franche

One of the more genuinely participatory technical bodies. If you can demonstrate a hazard category the benchmark misses, the working group is a place where that can actually change an artefact that companies are graded on. Requires technical credibility and sustained attendance.

Préoccupations couvertes

Product-level hazard and reliability: violent crime facilitation, CSAM, hate, self-harm, privacy, IP, defamation, specialised advice. Near-term deployment harms and misuse — NOT loss of control.

Canal gouvernemental

Moderate — AILuminate is cited in policy discussion of evaluation standards and MLCommons engages with NIST and the safety-institute network, but it has no formal mandate.

Sources

Une erreur ?