Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.¶
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., et al. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Intervention-Coupled Harm
- In AI safety, alignment training that reduces harmful outputs does so via the same gradients that produce sycophancy and over-refusal, and safety filters block prompts via the same classifier that miscategorises legitimate ones.
This sourceDocuments the helpfulness-harmlessness tension — safety training that reduces harmful outputs via the same RLHF objective that degrades helpfulness and drives over-refusal — the alignment tax as a coupled cost.
- In AI safety, alignment training that reduces harmful outputs does so via the same gradients that produce sycophancy and over-refusal, and safety filters block prompts via the same classifier that miscategorises legitimate ones.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:a00779131b56 · see in the full table