Français

Lieux à contacter

Anthropic

Type
AI company
Lieu
États-Unis — USA
Dernière vérification
2026-09-22
Prochaine vérification
2027-03-21

Pas encore traduit — affiché en anglais.

Comment les joindre

Ce qu'il fait

Safeguards team enforces the Usage Policy and runs Constitutional Classifiers; Alignment Science and Frontier Red Team run capability/misalignment evaluations feeding the Responsible Scaling Policy, overseen by a Responsible Scaling Officer.

Évaluation franche

Best-documented pathway of any lab. AI Frontiers' 2026 review of jailbreak disclosure found Anthropic runs two programs — a restricted NDA-bound CBRN bounty and an open cyber jailbreak program where 'anyone can submit without an NDA, findings may be publicly disclosed once timing is coordinated' — and named it as one of the few labs with a genuinely working route. Independent researchers pursuing cross-model universal jailbreaks in 2025-26 reported labs have 'very heterogeneous disclosure pathways (or none at all)' but that some, including Anthropic, remediated. No public evidence of usersafety@ being a black hole, but also little public evidence of reply rates.

Comment déposer

Security bugs: submit through Anthropic's HackerOne program with vulnerability type, technical detail, reproduction steps, URLs, PoC and impact. Model behaviour and jailbreaks: email usersafety@anthropic.com. This is the cleanest published example of the procedural split — a model-behaviour finding sent to a security intake will be closed as out of scope, and a security bug emailed to a safety address loses bounty eligibility.

Format

Security: standard vulnerability report. Model safety: prose email with exact prompts, transcripts, model version, date and reproducibility notes.

Calendrier

Rolling. Anthropic commits to acknowledge HackerOne submissions within three business days.

Ce qui se passe ensuite

Security reports are triaged on HackerOne with possible bounty. Model-safety emails have no published SLA, triage process or feedback loop.

Ce qu'il accepte

Two distinct lanes, and the distinction is the whole point. Security lane (via HackerOne): misconfigurations, CSRF, privilege escalation, SQL injection, XSS, directory traversal on internet-facing systems. Model-safety lane (via email): the policy states explicitly 'We welcome reports concerning safety issues, "jailbreaks," and similar concerns so that we can enhance the safety and harmlessness of our models' — these go to usersafety@anthropic.com.

Ce qu'il n'accepte pas

SSL/TLS issues without a proof of concept, physical intrusion, rate limiting on unauthenticated endpoints, social engineering, phishing, DoS, and widely publicised zero-days with no patch or a patch under 30 days old.

Géré par

Anthropic

Notes

Write to usersafety@anthropic.com with a reproducible transcript. Safe harbour is real but conditional: 'If you, in our sole determination, make a good faith effort to research and disclose vulnerabilities in accordance with this Policy... we will not pursue any legal action because of your research or responsible disclosure' — note 'in our sole determination'. Whistleblowing is internal only: the RSP Noncompliance Reporting and Anti-Retaliation Policy routes through NAVEX to the Responsible Scaling Officer and is written for employees and Board members, with no public-facing channel.

Sources

Une erreur ?