Human Compatible¶
Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control.
Cited by¶
3 citations across 3 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Agency
- In artificial systems and reinforcement learning, it is formalized as reward, world-model, and policy, and recent debates about constraint and alignment are debates about what kinds of agency such systems develop.
This sourceFrames AI agents in terms of objectives, world-models, and action selection and the alignment problem; supports reward-gaming, deception, and emergent sub-goal debates as one framework. (See also Sutton & Barto for the reward-model-policy formalism.)
- In artificial systems and reinforcement learning, it is formalized as reward, world-model, and policy, and recent debates about constraint and alignment are debates about what kinds of agency such systems develop.
- Homeostasis
- psychology and behavior (McEwen (1998) in NEJM developed allostasis as active regulation of physiological and behavioral variables including the stress response; emotion-regulation as homeostatic maintenance of affect), social systems (institutions that dampen disruptive forces to preserve norms — though this usage is often metaphorical),
This sourceViking. Argues that beneficial AI design requires homeostatic-style agents whose objectives reference human preferences as regulated variables; relevant for autoscaling, rate-limiting, and self-healing infrastructure framed as artificial homeostatic systems.
- psychology and behavior (McEwen (1998) in NEJM developed allostasis as active regulation of physiological and behavioral variables including the stress response; emotion-regulation as homeostatic maintenance of affect), social systems (institutions that dampen disruptive forces to preserve norms — though this usage is often metaphorical),
- Teleology
- AI alignment and intentional-systems theory Attributing goals to systems (agents, firms, algorithms) involves teleological framing; specifying reward functions and objective alignment, as Russell (2019) frames the alignment problem, requires identifying the system's "purpose" and managing uncertainty over the human preferences that purpose is meant to serve.
This sourceViking. Argues that beneficial AI design requires homeostatic-style agents whose objectives reference human preferences as regulated variables; relevant for autoscaling, rate-limiting, and self-healing infrastructure framed as artificial homeostatic systems.
- AI alignment and intentional-systems theory Attributing goals to systems (agents, firms, algorithms) involves teleological framing; specifying reward functions and objective alignment, as Russell (2019) frames the alignment problem, requires identifying the system's "purpose" and managing uncertainty over the human preferences that purpose is meant to serve.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:1d3f744aa10c · see in the full table