English

Organisations

METR (Model Evaluation & Threat Research)

Type
Civil society organisation
Place
United States — United States (Berkeley)
Last checked
2026-09-22
Next check due
2027-03-21

Ways to reach them

  • General enquiries · Email address
    Verified · 2026-09-22

    Anyone

  • Press · Email address
    Verified · 2026-09-22

    Media

  • Careers · Web form
    Verified · 2026-09-22

    Researchers and engineers

    METR also runs paid contractor/baseliner work for task development, which is an unusual paid entry route for technical outsiders.

  • Research updates · Web form
    Verified · 2026-09-22

    Anyone

  • Homepage · Homepage
    Unchecked

What it does

Research non-profit that evaluates frontier models for autonomous and dangerous capabilities and publishes the results. Best known for the time-horizon metric (length of task an agent can complete at 50% reliability), pre-deployment evaluations for OpenAI and Anthropic, and work on threats to evaluation integrity — generalised reward hacking and sandbagging. Formerly ARC Evals.

Honest assessment

Credible and influential well beyond its size, but heavily capacity-constrained and bound by NDAs with the labs it evaluates. A general inbox, no tip line. It will not act as a whistleblowing channel.

Concerns it covers

Dangerous autonomous capability and loss of control: autonomous replication, rogue deployment inside a lab, automated AI R&D acceleration, sabotage, and the meta-problem that models may deliberately underperform on evaluations.

Government channel

Strong — METR's evaluations are referenced in frontier-lab safety frameworks and by the US and UK safety/security institutes; its time-horizon results are widely cited in policy.

Sources

Something wrong here?