AI Safety via Debate¶
Irving, G., Christiano, & Amodei, D. (2018). AI Safety via Debate.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Dialectic
- Conversational AI structured as dialectical exchanges; debate-based alignment methods (Irving-Christiano-Amodei 2018) that use adversarial dialogue to surface reasoning; multi-agent dialogue for fact-checking and collaborative reasoning
This sourceProposes training agents via a two-player zero-sum debate game judged by a human, using adversarial dialogue to surface considerations and detect flaws; supports the debate-based-alignment example (191) and the AI-dialogue tension (T6).
- Conversational AI structured as dialectical exchanges; debate-based alignment methods (Irving-Christiano-Amodei 2018) that use adversarial dialogue to surface reasoning; multi-agent dialogue for fact-checking and collaborative reasoning
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:bfc4f178d347 · see in the full table