Deep reinforcement learning from human preferences¶
Christiano, P. F., Leike, Brown, Martic, Legg, & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems.
Cited by¶
2 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Preference
- Utility functions, rankings, revealed choices, policy priorities, qualitative value orderings, and learned reward signals are all implementations of the same ordering relation; the relation is the prime, the implementation is local technology — a substrate range that runs from Debreu's (1954) representation theorem for continuous preference orderings to Christiano et al.'s (2017) deep-RL reward models fit from pairwise human comparisons.
This sourceIntroduces the now-standard pipeline of fitting a reward model from human pairwise preference comparisons over agent trajectories and optimizing a policy against it; canonical reference for preference learning as a substrate for deep RL.
- Utility functions, rankings, revealed choices, policy priorities, qualitative value orderings, and learned reward signals are all implementations of the same ordering relation; the relation is the prime, the implementation is local technology — a substrate range that runs from Debreu's (1954) representation theorem for continuous preference orderings to Christiano et al.'s (2017) deep-RL reward models fit from pairwise human comparisons.
- Refinement
- Machine learning & neural networks: Gradient descent as refinement of weights toward lower loss, backpropagation as feedback signal, hyperparameter tuning, RLHF (reinforcement learning from human feedback) as refinement of language model outputs toward human-preferred behaviors, the framework Christiano et al. (2017) introduced for aligning agents with human preferences.
This sourceIntroduces RLHF: a learned reward model trained on human pairwise preferences guides iterative refinement of agent policy toward human-aligned behavior.
- Machine learning & neural networks: Gradient descent as refinement of weights toward lower loss, backpropagation as feedback signal, hyperparameter tuning, RLHF (reinforcement learning from human feedback) as refinement of language model outputs toward human-preferred behaviors, the framework Christiano et al. (2017) introduced for aligning agents with human preferences.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:9ffcead8b0da · see in the full table