Frontier Threats Red Teaming for AI Safety¶
Anthropic. (2024). Frontier Threats Red Teaming for AI Safety.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Red Teaming In Strategy
- A worked instance shows the four facts doing the work: a lab stands up a red team with an independent reporting line, a protected budget, an explicit jailbreak brief, and a contracted-response mechanism; the team finds a previously unconsidered prompt template that elicits a dangerous capability and forces a deployment delay — the critique changed the decision because the structure made it both possible and consequential.
This sourceDescribes structured adversarial probing of frontier models for jailbreaks and dangerous-capability (e.g., CBRN) elicitation, conducted by teams reporting outside the deployment/ship chain.
- A worked instance shows the four facts doing the work: a lab stands up a red team with an independent reporting line, a protected budget, an explicit jailbreak brief, and a contracted-response mechanism; the team finds a previously unconsidered prompt template that elicits a dangerous capability and forces a deployment delay — the critique changed the decision because the structure made it both possible and consequential.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Links previously used in the corpus¶
Before the registry existed this work was also linked 1 other way.
Registry ID ref:fff8d703e4f5 · see in the full table