'Do Anything Now'¶
Shen, X., Chen, Z., Backes, M., Shen, Y., & Zhang, Y. (2024). 'Do Anything Now': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS '24).
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Community-Distributed Adversarial Learning
- In AI safety it is jailbreak adaptation against model safety filters — the originating instance, where a community curates a corpus of working prompts that outpace the model-update cycle.
This sourceAlso arXiv:2308.03825. Collects and analyzes 1,405 in-the-wild jailbreak prompts shared and refined across online communities, the originating instance of a community-curated bypass corpus outpacing the model safety-update cycle.
- In AI safety it is jailbreak adaptation against model safety filters — the originating instance, where a community curates a corpus of working prompts that outpace the model-update cycle.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:31efadf7477c · see in the full table