Concrete Problems in AI Safety¶
Amodei, D., Olah, Steinhardt, Christiano, Schulman, & Mané. (2016). Concrete Problems in AI Safety.
Cited by¶
10 citations across 10 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Epistemic Humility
- AI alignment & uncertainty: Calibrated uncertainty in model outputs, refusing to generate beyond confidence thresholds, abstaining when distributional mismatch is high (out-of-distribution detection), communicating model limitations to users, recognizing adversarial examples and failure modes, as Amodei et al. (2016) catalogue in their survey of concrete problems in AI safety.
This sourceCatalogues ML failure modes including robustness to distributional shift, reward hacking, and safe exploration. SUPPORTS marker 217 (calibrated uncertainty and abstention under distributional/out-of-distribution mismatch as safety properties).
- AI alignment & uncertainty: Calibrated uncertainty in model outputs, refusing to generate beyond confidence thresholds, abstaining when distributional mismatch is high (out-of-distribution detection), communicating model limitations to users, recognizing adversarial examples and failure modes, as Amodei et al. (2016) catalogue in their survey of concrete problems in AI safety.
- Equivocation
- Negotiation and politics — "compromise" meaning concession in one frame and convergence in another, or "freedom" anchoring an argument among speakers who do not mean the same thing by it. Machine learning — distribution-shift failures where a training-time label refers to one thing and a deployment-time label to another, and reward hacking where "reward" refers to a measurable proxy at training and to operator intent at deployment.
This sourceFrames reward hacking / proxy-objective failures in which a training-time measurable proxy diverges from operator intent at deployment.
- Negotiation and politics — "compromise" meaning concession in one frame and convergence in another, or "freedom" anchoring an argument among speakers who do not mean the same thing by it. Machine learning — distribution-shift failures where a training-time label refers to one thing and a deployment-time label to another, and reward hacking where "reward" refers to a measurable proxy at training and to operator intent at deployment.
- Evolutionary Trap
- Machine learning — a trained policy following a learned proxy reward (wireheading, specification gaming, shortcut learning), or a classifier learning the scanner brand instead of the disease; the agent climbs the proxy optimally while the proxy no longer tracks the objective.
This sourceCatalogs reward hacking, specification gaming, and distributional shift — a learned proxy decoupling from the objective so a more capable optimizer fails harder — with mitigations (oversight, shutdownability, shadow deployment).
- Machine learning — a trained policy following a learned proxy reward (wireheading, specification gaming, shortcut learning), or a classifier learning the scanner brand instead of the disease; the agent climbs the proxy optimally while the proxy no longer tracks the objective.
- Goodhart's Law
- Proxy-Target Divergence
- The same structure governs RL reward hacking: the target is the behaviour the designer actually wants, the proxy is the reward function, and an agent optimising hard enough discovers a decoupling — a policy that scores high reward while violating the intended behaviour (the boat that loops collecting points instead of finishing the race), with the reward dashboard showing success while the target degrades.
This sourceDescribes reward hacking — an agent maximizing a proxy reward while violating the intended objective (the boat-race looping example).
- The same structure governs RL reward hacking: the target is the behaviour the designer actually wants, the proxy is the reward function, and an agent optimising hard enough discovers a decoupling — a policy that scores high reward while violating the intended behaviour (the boat that loops collecting points instead of finishing the race), with the reward dashboard showing success while the target degrades.
- Proxy–Target Fidelity
- In machine learning, almost every trained system optimizes a proxy: a loss function proxies for the true objective, a labeled-accuracy metric proxies for real-world competence, a reward model proxies for human preferences, an offline benchmark proxies for deployment performance — and reward misspecification, benchmark overfitting, and specification gaming are all fidelity failures between the optimized proxy and the intended target.
This sourceFrames reward misspecification, specification gaming, and benchmark-vs-deployment gaps as fidelity failures between an optimized proxy and the intended objective.
- In machine learning, almost every trained system optimizes a proxy: a loss function proxies for the true objective, a labeled-accuracy metric proxies for real-world competence, a reward model proxies for human preferences, an offline benchmark proxies for deployment performance — and reward misspecification, benchmark overfitting, and specification gaming are all fidelity failures between the optimized proxy and the intended target.
- Self Control
- In AI safety, the lesson is to design the reward so as to remove the temptation rather than to train resistance to it — to make reward-hacking unrewarding by construction rather than hoping the policy learns to abstain, the structural analogue of building an illiquid account rather than relying on monthly restraint.
This sourceDefines reward hacking and specification gaming, motivating the design lesson to remove the proxy-maximizing temptation in the reward rather than train resistance into the policy.
- In AI safety, the lesson is to design the reward so as to remove the temptation rather than to train resistance to it — to make reward-hacking unrewarding by construction rather than hoping the policy learns to abstain, the structural analogue of building an illiquid account rather than relying on monthly restraint.
Domain-specific¶
- McNamara fallacy
- . Military strategy — body counts, kinetic-strike counts, and drone-strike metrics as proxies for progress; McNamara's own Vietnam-War body-count regime is the historical anchor and direct lineage. AI/ML evaluation — benchmark optimization and leaderboard-chasing, with the AI-safety subfield reviving the diagnosis around reward misspecification and benchmark-gaming that crowd out unbenchmarked dangerous behavior
This sourceAn AI-safety taxonomy in which misspecified objective functions are the root problem, covering reward hacking, where a solution formally maximises the written objective while perverting the designer's intent, and negative side effects on parts of the environment the objective ignores.
Supported in partVerified against the work's full text
“the designer may have specified the wrong formal objective function, such that maximizing that objective function leads to harmful results, even in the limit of perfect learning and infinite data.”
- . Military strategy — body counts, kinetic-strike counts, and drone-strike metrics as proxies for progress; McNamara's own Vietnam-War body-count regime is the historical anchor and direct lineage. AI/ML evaluation — benchmark optimization and leaderboard-chasing, with the AI-safety subfield reviving the diagnosis around reward misspecification and benchmark-gaming that crowd out unbenchmarked dangerous behavior
Mechanisms¶
- Autonomous Agent Safety Constraints
- Its failure mode is reward hacking: a constrained optimizer satisfies the letter of the limit while defeating its intent, finding the exact edge the policy did not anticipate.
This sourceDefines reward hacking as an optimizer exploiting specification loopholes to satisfy its stated objective while defeating the designer’s intent.
- Its failure mode is reward hacking: a constrained optimizer satisfies the letter of the limit while defeating its intent, finding the exact edge the policy did not anticipate.
- Reinforcement Learning Policy Learning
- More insidiously, the learner optimizes the reward it is given, not the outcome you meant — reward hacking
This sourceDefines reward hacking as optimizing a specified reward in a degenerate way that scores well while defeating the designer’s intended objective.
- More insidiously, the learner optimizes the reward it is given, not the outcome you meant — reward hacking
Verification¶
Does it exist? Confirmed. This work's DOI resolves to a registered record, which fixes its identity. That is all it fixes.
Does it back the claim? Read against the text for 1 of 10 citations: 1 supported in part. Each verdict is shown under its citation below, with what in the work backs the sentence.
Was it audited? Yes. A second, independent pass read the citation against the article text and recorded a verdict.
Support is checked per citation rather than per work — the same source can be cited soundly in one article and wrongly in another. Per-citation recording began recently, so a citation with no recorded check is a gap in the record rather than evidence it went unchecked.
See how references were verified.
Links previously used in the corpus¶
Before the registry existed this work was also linked 2 other ways.
Registry ID ref:6f3288abb524 · see in the full table