Français

Lieux à contacter

Anthropic Model Safety Bug Bounty

Type
Public reporting channel
Lieu
International — Global
Dernière vérification
2026-09-22
Prochaine vérification
2027-03-21

Pas encore traduit — affiché en anglais.

Comment les joindre

Ce qu'il fait

Universal jailbreaks that defeat Anthropic's Constitutional Classifiers — generalised techniques that work across many prompts, tested against a supplied set of harmful questions concerning CBRN/biological threat information. Narrow, context-specific exploits do not qualify.

Évaluation franche

Pays real money and is a genuine channel to the safety team, but it is gated: you must be accepted, sign an NDA, and you may not publish the jailbreak, the test questions or classifier details without written consent. You may disclose that you participate. Not a route for someone who wants to speak publicly, and not open-entry.

Comment déposer

Complete a Google Form application and create a HackerOne account first. If accepted you receive a HackerOne invitation and must create a Claude Console account using your @wearehackerone.com alias. An NDA is mandatory. Applications reviewed on a rolling basis.

Format

HackerOne report with detailed methodology; rewards graded on a sliding scale against internal criteria for response detail and accuracy.

Calendrier

Rolling applications; invite-only participation.

Ce qui se passe ensuite

Graded against internal criteria; up to $35,000 per novel universal jailbreak. Accepted participants get free access to a model alias for authorised red-teaming only.

Ce qu'il accepte

Universal jailbreaks that defeat Anthropic's Constitutional Classifiers — generalised techniques that work across many prompts, tested against a supplied set of harmful questions concerning CBRN/biological threat information. Narrow, context-specific exploits do not qualify.

Ce qu'il n'accepte pas

Infrastructure vulnerabilities (misconfigurations, CSRF, privilege escalation, SQLi, XSS, directory traversal) — those belong under the Responsible Disclosure Policy. Also excludes single-prompt or narrow jailbreaks.

Géré par

Anthropic, run on HackerOne

Sources

Une erreur ?