English

Places to contact

Anthropic Model Safety Bug Bounty

Type
Public reporting channel
Place
International — Global
Last checked
2026-09-22
Next check due
2027-03-21

Ways to reach them

What it does

Universal jailbreaks that defeat Anthropic's Constitutional Classifiers — generalised techniques that work across many prompts, tested against a supplied set of harmful questions concerning CBRN/biological threat information. Narrow, context-specific exploits do not qualify.

Honest assessment

Pays real money and is a genuine channel to the safety team, but it is gated: you must be accepted, sign an NDA, and you may not publish the jailbreak, the test questions or classifier details without written consent. You may disclose that you participate. Not a route for someone who wants to speak publicly, and not open-entry.

How to file

Complete a Google Form application and create a HackerOne account first. If accepted you receive a HackerOne invitation and must create a Claude Console account using your @wearehackerone.com alias. An NDA is mandatory. Applications reviewed on a rolling basis.

Format

HackerOne report with detailed methodology; rewards graded on a sliding scale against internal criteria for response detail and accuracy.

Timing

Rolling applications; invite-only participation.

What happens after

Graded against internal criteria; up to $35,000 per novel universal jailbreak. Accepted participants get free access to a model alias for authorised red-teaming only.

What it accepts

Universal jailbreaks that defeat Anthropic's Constitutional Classifiers — generalised techniques that work across many prompts, tested against a supplied set of harmful questions concerning CBRN/biological threat information. Narrow, context-specific exploits do not qualify.

What it does not accept

Infrastructure vulnerabilities (misconfigurations, CSRF, privilege escalation, SQLi, XSS, directory traversal) — those belong under the Responsible Disclosure Policy. Also excludes single-prompt or narrow jailbreaks.

Operated by

Anthropic, run on HackerOne

Sources

Something wrong here?