English

Bodies with power

MLCommons — AILuminate benchmark (AI Risk & Reliability working group)

Type
Standards body
Place
International — International (engineering consortium)
Last checked
2026-09-22
Next check due
2027-03-21

Ways to reach them

  • AI Risk & Reliability working group — open to participants · Official membership list
    Verification expired · 2026-09-22

    This route is retained as reference only. Its verification is missing, more than 180 days old, or the contact window is not currently open; do not rely on it until it is re-checked.

    Researchers and practitioners; working-group participation is the substantive route

    MLCommons working groups run open calls and mailing lists; this is a genuine build-it-with-us channel.

  • Get involved · Official membership list
    Verification expired · 2026-09-22

    This route is retained as reference only. Its verification is missing, more than 180 days old, or the contact window is not currently open; do not rely on it until it is re-checked.

    Anyone

  • GitHub — code, taxonomy and prompt-set contributions/issues · Written submission
    Verified · 2026-09-22

    Anyone

  • MLCommons Discord · Web form
    Verified · 2026-09-22

    Anyone

  • Public benchmark results · Web form
    Verified · 2026-09-22

    Read-only

  • Homepage · Homepage
    Unchecked

What it does

AILuminate is a standardised AI risk assessment benchmark for chat and vision-language models across 12 hazard categories, built with industry, academia and civil society; at last check 59,624 test prompts, 477 test images and 109 models benchmarked, with expanding agentic, multimodal and jailbreak/security suites. Produces public grades per model.

Honest assessment

One of the more genuinely participatory technical bodies. If you can demonstrate a hazard category the benchmark misses, the working group is a place where that can actually change an artefact that companies are graded on. Requires technical credibility and sustained attendance.

Concerns it covers

Product-level hazard and reliability: violent crime facilitation, CSAM, hate, self-harm, privacy, IP, defamation, specialised advice. Near-term deployment harms and misuse — NOT loss of control.

Government channel

Moderate — AILuminate is cited in policy discussion of evaluation standards and MLCommons engages with NIST and the safety-institute network, but it has no formal mandate.

Sources

Something wrong here?