METR (Model Evaluation & Threat Research)
- Type
- Civil society organisation
- Lieu
- États-Unis — United States (Berkeley)
- Dernière vérification
- 2026-09-22
- Prochaine vérification
- 2027-03-21
Pas encore traduit — affiché en anglais.
Comment les joindre
- General enquiries
- Press
- Careers
- Research updates
- Homepage
Ce qu'il fait
Research non-profit that evaluates frontier models for autonomous and dangerous capabilities and publishes the results. Best known for the time-horizon metric (length of task an agent can complete at 50% reliability), pre-deployment evaluations for OpenAI and Anthropic, and work on threats to evaluation integrity — generalised reward hacking and sandbagging. Formerly ARC Evals.
Évaluation franche
Credible and influential well beyond its size, but heavily capacity-constrained and bound by NDAs with the labs it evaluates. A general inbox, no tip line. It will not act as a whistleblowing channel.
Préoccupations couvertes
Dangerous autonomous capability and loss of control: autonomous replication, rogue deployment inside a lab, automated AI R&D acceleration, sabotage, and the meta-problem that models may deliberately underperform on evaluations.
Canal gouvernemental
Strong — METR's evaluations are referenced in frontier-lab safety frameworks and by the US and UK safety/security institutes; its time-horizon results are widely cited in policy.