Redwood Research
- Type
- Civil society organisation
- Lieu
- États-Unis — United States (Berkeley)
- Dernière vérification
- 2026-09-22
- Prochaine vérification
- 2027-03-21
Pas encore traduit — affiché en anglais.
Comment les joindre
- General enquiries
- Research engagement via public writing
- Homepage
Ce qu'il fait
Non-profit research organisation that effectively founded the 'AI control' agenda: designing and red-teaming protocols that keep deployments safe even if the model is actively trying to subvert them. Known for control evaluations, monitoring protocols for malign LLM agents, the alignment-faking work with Anthropic, and helping developers build safety cases.
Évaluation franche
Small and technically serious; will engage substantively with a good technical argument, and its researchers are notably responsive in public research forums. Will not engage with policy complaints or non-technical concerns.
Préoccupations couvertes
Intentional misalignment and scheming — models that purposefully act against their developers' interests. Distinctively assumes alignment may fail and asks what safety measures survive that assumption. Loss-of-control, not misuse or near-term harms.
Canal gouvernemental
Indirect — AI control has been adopted into frontier-lab safety frameworks and cited in safety-institute work; Redwood advises labs rather than governments.