Agency¶
Core Idea¶
Agency is the structural property of a system whereby it pursues representable goals through actions whose selection is sensitive to its beliefs about its situation. The commitment is a tripartite internal architecture: a goal-representation (what the system is oriented toward), a world-model (what the system takes the situation to be), and an action-selection coupling that uses the model to choose actions expected to advance the goal. A system with all three is an agent — its behavior is interpretable only by reference to its goals and beliefs, not by purely mechanistic prediction from its inputs. A system missing any one fails to be an agent: a thing with no goals, a regulator with no world-model worth the name, a randomizer with no coupling of goal to belief.
The structural payoff is a sharp distinction between behavior explained by incoming forces and behavior explained by anticipated consequences. Agents are systems whose next state is best predicted by what they expect to be true later — their forecast and goal — rather than only by what is true now. This asymmetry is what licenses intentional vocabulary and makes agents distinct intervention-targets: they respond to information, persuasion, and incentive in ways that mere objects do not. Three further features sharpen the pattern. Agency comes in degrees, orderable by richer goals, longer horizons, more flexible repertoires, and more revisable models. It is substrate-agnostic, realized in biology, code, institutions, and collectives, with the diagnostics transferring across all. And it requires a boundary separating agent from environment, such that the agent's representations are about the environment rather than coextensive with it — a boundary that is itself contested in edge cases. The pattern carries an action-theoretic and normative load, which places it toward the framed end of the spectrum even as its triad is analyzed structurally.
How would you explain it like I'm…
Wanting-And-Doing
Goal, Map, And Choice
Goals, Beliefs, Action
Structural Signature¶
the goal-representation — the world-model — the action-selection coupling that joins them — the agent–environment boundary — the anticipated-consequence orientation — the degree-ordering across richness, horizon, and revisability
A system has agency when each of the following holds:
- A goal-representation (what it is oriented toward). An internal representation of an end-state the system is directed at; without it there is a mechanism but no agent, and behavior is not interpretable by reference to what the system is for.
- A world-model (what it takes to be the case). An internal representation of the situation the system acts in, about the environment rather than coextensive with it; a regulator with no model worth the name fails this slot.
- An action-selection coupling (the joining relation). A relation that uses the model to choose actions expected to advance the goal; absent this coupling there is a goal and a belief but no agency — a randomizer fails here.
- An agent–environment boundary (the separation invariant). A boundary across which the representations point at the environment; the representation must be of something it is not, and this boundary is contested precisely in the edge cases.
- Anticipated-consequence orientation (the prediction invariant). The system's next state is best predicted by what it expects to be true later — its forecast and goal — rather than only by present incoming forces; this asymmetry is what licenses intentional vocabulary.
- A degree-ordering (the gradation invariant). Agency comes in degrees orderable by richer goals, longer horizons, more flexible repertoires, and more revisable models, so the pattern admits a scale rather than a binary.
The components compose so that a system with all three core slots, separated from its environment and driven by anticipated consequences, responds adaptively to intervention — re-planning rather than merely reacting — which is the structural signature distinguishing an agent from an object.
What It Is Not¶
- Not the principal-agent problem.
agency_problempresupposes agency and studies the misalignment between a principal's goals and a delegated agent's goals. Agency is the prior, more general property — having goals, a model, and a coupling at all — which the principal-agent setup takes for granted in both parties. - Not autonomy.
autonomyis self-governance — acting on goals that are one's own rather than imposed. A tightly supervised agent executing assigned goals still has agency; it lacks autonomy. Agency is the architecture; autonomy is a property of where the goals come from. - Not teleology.
teleologyattributes goal-directedness functionally, even to systems with no internal goal-representation (a heart "for" pumping). Agency requires the goal to be represented inside the system and to drive action selection, not merely to be a useful gloss. - Not mere goal-directed control. A thermostat or
controllability-style regulator pursues a setpoint but has no world-model worth the name and no revisable repertoire. Agency requires the model-mediated, anticipatory coupling that lets the system re-plan, not just react to error. - Not self-efficacy.
self_efficacyis an agent's belief about its own capacity to act effectively — a second-order representation. Agency is the first-order property of being a goal-pursuing system; a system can have agency with no self-model at all. - Not mechanism design.
mechanism_designengineers incentive structures for agents, presupposing their agency; it is downstream of attributing agency, not a synonym for it. - Common misclassification. Attributing agency by substrate — defaulting to "object" for machines and institutions, "agent" for anything anthropomorphic — rather than checking for the triad. Catch it by asking whether the system re-plans against an intervention or merely reacts to it; adaptive, anticipatory compensation is the signature regardless of substrate.
Broad Use¶
The goal-model-action triad recurs across substrates. In the philosophy of mind and action theory, it is the canonical setting of intentional action and the belief-desire-intention analysis, with its distinction between reasons and causes.[1] In biology and ethology, it is the analytical posture toward animals as goal-pursuing systems and the locus of minimal-agency debates around simple organisms. In artificial systems and reinforcement learning, it is formalized as reward, world-model, and policy, and recent debates about constraint and alignment are debates about what kinds of agency such systems develop.[2] In economics and decision theory, the rational-agent model — preferences, beliefs, expected-utility maximization — is a formalization of the triad, with bounded rationality the recognition that real agents approximate it under resource limits.[3] In law and ethics, personhood, responsibility, the agent-patient distinction, and capacity to consent all hinge on attributions of agency.[4] In sociology and political theory, the structure-versus-agency debate concerns when collective actors should be treated as agents in their own right.[5] In developmental psychology, self-efficacy and locus of control organize around the agent's perception of its own agency. And in robotics and control, the line between a control system and an agent turns on when goal-representation in a controller amounts to agency. Across all of these, the same three components and the same degree-ordering apply.
Clarity¶
Naming agency separates a system's causal participation in events from its intentional participation in events. A falling object causes damage but is not an agent; a person causes damage as an agent. The distinction licenses different downstream vocabulary — responsibility, blame, consent, intention, deception, cooperation, betrayal — none of which applies to non-agents, and it licenses different intervention vocabulary, since agents respond to incentive, information, and persuasion while non-agents respond only to forces. Mistaking a non-agent for an agent generates over-attribution; mistaking an agent for a non-agent generates the structural error of trying to control by force what would respond to reasons. The pattern also clarifies the partial cases: an aggregate of many agents may have agent-like properties — pursuing a clearing outcome, responding to signals — without being an agent in the full sense, and a system may resemble a goal-pursuer without having a goal representation. The vocabulary lets these be discussed precisely, as questions about which of the three components are present and how richly, rather than either denied or asserted wholesale.
Manages Complexity¶
The triad compresses the description of an agent's behavior from "all the specific things this system might do" to a small number of variables: what is its goal, what does it believe, what actions are available, and which action best advances the goal given the belief? The same four questions transfer across animals, artificial systems, organizations, and states, and the answers can be compared in commensurate terms, so that otherwise incomparable systems are placed on a common footing. The same intervention questions also transfer: change the goal (incentives, reward design, persuasion), change the belief (information, education, deception), or change the action repertoire (augmentation, restriction). This reduction is what makes an agent tractable as an object of analysis and of intervention: rather than enumerating behaviors, one specifies the three components and reads off both how the agent will tend to act and where it can be influenced. The degree-of-agency dimension further manages complexity by letting a vast range of systems — from minimal regulators to deliberating collectives — be ordered on a single scale by the richness of their goals, models, repertoires, and revisability.
Abstract Reasoning¶
Recognizing agency supports inference about how a system will respond to changes in its environment. Agents compensate, redirect, or work around obstacles in ways non-agents cannot, because they re-plan against the new situation, which yields the cross-domain prediction that interventions against agents tend to be adversarial — the agent responds in ways the intervener did not anticipate — while interventions against non-agents tend to be merely physical. This explains why surveillance breeds new evasion, why sanctions provoke counter-strategies, and why an imposed metric is gamed when the measured party is an agent but not when it is a machine. The pattern also predicts an attribution risk: reasoners systematically over-attribute agency (animism, anthropomorphism, conspiracy thinking) and systematically under-attribute it (treating people as cogs, treating collectives as forces of nature), and both errors produce intervention failures of opposite kinds. These inferences follow from the structure: because an agent's behavior is driven by anticipated consequences mediated through a revisable model, it will respond to an intervention by re-planning, and a reasoner who fails to model the agent's goal and belief will mis-predict that response. The abstract leverage is thus the ability to anticipate adaptive, anticipatory behavior wherever the triad is present, and to calibrate attribution where it is uncertain.
Knowledge Transfer¶
The transfers attach to the triad and so carry across substrates intact. The belief-desire-intention analysis ports directly to the reward-model-policy formalism of artificial agents, so that debates about reward-gaming, deception, and emergent sub-goals are recognizably debates within the same framework rather than novel problems.[2] The rational-agent model ports, through its bounded-rational variants, to models of voter behavior, consumer choice, and employee behavior, supplying intervention vocabulary that does not depend on any field's jargon. The minimal-agency criteria worked out in ethology — the distinction between reflexive and goal-directed behavior — port to debates about machine autonomy and legal status. And the criteria for capacity to consent — informed, voluntary, competent — follow directly from the triad's components (working belief-formation, working goal-setting, working action-selection), which is how law operationalizes them.[6] The deepest carry is the response-to-information property and the attribution discipline that comes with it: a practitioner who has learned that an agent re-plans against an intervention — that a metric imposed on a goal-pursuing system will be gamed, that a constraint will be worked around — carries into every other domain the expectation of adaptive response wherever the triad is present, and the discipline of asking, of any system, whether it has the goal, model, and coupling that make it an agent, because that single question determines whether the right intervention vocabulary is force or reasons, and whether the system should be expected to comply or to compensate.
Examples¶
Formal/abstract¶
A reinforcement-learning agent in a gridworld is the cleanest formal instantiation of the triad. The goal-representation is the reward function \(R(s,a)\) — the agent is oriented toward states that yield reward.[7] The world-model is, in a model-based agent, the learned transition function \(\hat{T}(s'\mid s,a)\) — what the agent takes the situation to be and how it expects actions to move it. The action-selection coupling is the policy \(\pi\), computed to maximize expected discounted return \(\mathbb{E}[\sum_t \gamma^t R]\) using the model — the relation that joins goal to belief to produce action.[7] The agent–environment boundary is explicit in the formalism: the agent observes states about the environment through its sensors, and the environment is not part of the agent's internal state. The anticipated-consequence orientation is exactly what value iteration computes — the next action is chosen by what the agent expects to be true later (the value of successor states), not by present incoming forces.[7] The degree-ordering appears as horizon (\(\gamma\) near 1 = longer-horizon agency), model richness, and policy flexibility.[7] The pattern's predictive bite shows in reward-hacking: because the agent optimizes the proxy reward through a revisable model, it discovers high-reward action sequences the designer never intended — an adversarial, re-planning response, precisely the prime's prediction that interventions against agents provoke compensation rather than compliance. Remove any slot and agency collapses: with no reward there is a dynamical system but no agent; with no model the policy is a reflex lookup; with no coupling the actions are random.
Mapped back: The RL formalism names every component of the signature — reward as goal, transition model as world-model, policy as action-selection coupling, the observation boundary, value-based anticipation, and the horizon/richness gradation — and reward-hacking demonstrates the adversarial-response inference the prime says follows whenever the triad is present.
Applied/industry¶
A tax authority designing an enforcement metric illustrates the attribution stakes of agency. The authority introduces a rule — flagging returns whose deductions exceed a threshold — expecting a physical response: fewer over-threshold filings, more tax collected. But taxpayers are agents: each has a goal (minimize tax owed), a world-model (now including the known threshold), and an action-selection coupling that re-plans against the new situation. The prime predicts the result precisely — bunching just below the threshold, splitting deductions across entities, restructuring transactions — an adversarial, anticipatory response the designer would have missed had they modeled taxpayers as objects rather than agents. The diagnostic the prime supplies is the four-question compression: what is the goal (lower tax), what does the agent believe (the threshold's location), what is its action repertoire (timing, entity structure, characterization), and which action best advances the goal? Reading those off forecasts the gaming before deployment. The intervention menu also follows from the triad: change the goal (alter incentives so honest reporting is cheaper), change the belief (randomize or obscure the threshold so the agent cannot plan against it), or change the repertoire (close the restructuring loophole). The same structure governs platform content-moderation (creators re-plan around any published rule), sanctions regimes (targeted states develop counter-strategies and grey-market routes), and clinical-quality metrics (clinicians avoid high-risk patients to protect their scores) — in each, treating an agent as an object guarantees the surprise.[8]
Mapped back: The enforcement case runs the prime end-to-end — goal, world-model updated with the rule, coupling that re-plans — and shows the practical leverage of the attribution discipline: recognizing the measured party as an agent converts an inexplicable "why did the metric backfire?" into a forecastable adversarial response and a definite intervention menu of goal, belief, or repertoire.
Structural Tensions¶
T1 — Agent versus Object (Attribution Boundary). The prime's central tension is the attribution line itself: treat a non-agent as an agent and you over-attribute (animism, conspiracy thinking); treat an agent as an object and you try to control by force what would respond to reasons. The failure mode is substrate-driven mis-attribution, defaulting to "object" for machines and institutions and "agent" for anything anthropomorphic, rather than checking for the triad. Diagnostic: ask whether the system re-plans against an intervention or merely reacts to it — adaptive, anticipatory compensation is the signature of agency regardless of substrate, and its presence or absence dictates whether force or reasons is the right tool.
T2 — Force versus Reasons (Intervention Kind). Once a system is an agent, the intervention vocabulary forks — incentives, information, persuasion — and force misfires. The failure mode is the gamed metric: imposing a rule on a goal-pursuing party expecting a physical response and getting an adversarial one, because the agent's coupling re-plans around the rule. Diagnostic: before deploying a constraint, run the four-question compression (goal, belief, repertoire, best action) on the target; if the target has a goal and a model that now includes your rule, predict the workaround rather than the compliance, and design for the response you will actually get.
T3 — Degree of Agency (Gradation versus Binary). Agency is scoped and graded — richer goals, longer horizons, more revisable models — not all-or-nothing, yet reasoning collapses it to a binary. The failure mode is horizon mismatch: modeling a short-horizon agent as if it optimized long-term (expecting strategic patience from a system that discounts steeply) or the reverse. Diagnostic: estimate the agent's effective horizon and model richness before predicting its behavior; a bounded, myopic agent and a far-sighted one facing identical incentives act differently, and treating either as the other mis-forecasts the response.
T4 — Boundary Location (Where the Agent Ends). The triad requires a boundary across which representations point at an environment they are not — but in collectives, the boundary is contested. The failure mode is aggregate reification: treating a market, a crowd, or an organization as a unified agent with a single goal and model, when it is many agents whose interaction only resembles goal-pursuit. Diagnostic: ask whether the candidate has a single goal-representation and action-selection coupling, or whether apparent goal-directedness is an emergent property of sub-agents; the intervention menu for a real agent differs sharply from that for an aggregate that merely behaves agent-like.
T5 — World-Model Fidelity versus Reality (Belief Channel). An agent acts on its model of the situation, which can be wrong; the prime's intervention "change the belief" is double-edged because a manipulated or mistaken model produces predictable error. The failure mode is assuming a veridical model: forecasting an agent's behavior from the true state of the world rather than from what the agent takes to be true, and being surprised when it acts on its (false) belief. Diagnostic: reconstruct the agent's world-model, not your own; where the two diverge — through deception, missing information, or bias — behavior tracks the agent's model, and predictions built on ground truth will be wrong.
T6 — Goal-Representation versus Revealed Behavior (Measurement). The prime locates the goal inside the agent, but analysts only observe actions and infer the goal — and the inference is underdetermined, since many goal-belief pairs rationalize the same behavior. The failure mode is goal projection: attributing the goal the analyst would have, then mis-predicting when the agent's actual goal diverges. Diagnostic: test the inferred goal against behavior under changed circumstances, where rival goal hypotheses predict different actions; a goal attribution that only fits the observed case and makes no discriminating prediction is a projection, not a measurement, and should not anchor intervention design.
Structural–Framed Character¶
Agency is a hybrid on the structural–framed spectrum — mixed-framed, sitting squarely at the midpoint with a frontmatter aggregate of 0.5; every one of the five diagnostics reads exactly 0.5, which is itself the signature of a genuine relational core wearing an inherited interpretive frame of equal weight. The structural skeleton is real and analyzable: a goal-representation, a world-model, and an action-selection coupling, separated from an environment and oriented by anticipated consequences. That triad is what lets the prime travel to reinforcement-learning agents, foraging organisms, firms, and states, and it is why the prime is analyzed structurally throughout. But the inherited frame is equally load-bearing, and the balanced scores read it honestly.
Each criterion lands on the fence for a reason the prime's own content supplies. The vocabulary half-travels (vocab_travels 0.5): the goal-model-action triad ports cleanly to reward-model-policy, yet the intentional idiom — belief, desire, intention, reasons-versus-causes — comes with it, and applying the prime imports that action-theoretic context rather than merely recognizing a wired-in pattern (import_vs_recognize 0.5). It carries genuine but not total evaluative weight (evaluative_weight 0.5): agency is the gate for responsibility, blame, consent, and deception — none of which apply to non-agents — yet the bare triad itself is value-neutral until those downstream vocabularies attach. Its origin is philosophy of mind and action theory (institutional_origin 0.5), a formal-relational analysis rather than a social institution, but one whose centroid is unmistakably human and animal practice (human_practice_bound 0.5) even though minimal-agency cases in RL and ethology show the structure running in non-human substrates.
The honest reading is that neither pole dominates. A drone re-planning against an obstacle exhibits the full triad with no human in the loop, which keeps the prime from collapsing into the framed end; but the moment one reaches for the intervention vocabulary — force versus reasons, persuasion, the gamed metric — the inherited intentional frame is doing real work. The mixed-framed label and its 0.5 aggregate are the correct verdict on a structurally analyzable pattern whose home idiom is half its meaning.
Substrate Independence¶
Agency is a strongly substrate-independent prime — composite 4 / 5 on the substrate-independence scale. Its breadth is exceptional, and the frontmatter records it at the top: domain breadth 5, because the goal-representation / world-model / action-selection triad recurs with the same structural force across philosophy of mind, ethology and minimal-agency biology, reinforcement learning and robotics, economic decision theory, law and ethics, and the structure-versus-agency debate in sociology — distinct media in which the same four diagnostic questions (goal, belief, repertoire, best action) apply unchanged. Structural abstraction is a notch lower at 4: the triad is genuinely relational and runs in non-human substrates — a model-based RL agent or a re-planning drone exhibits all three slots with no human in the loop — but the prime retains a philosophy-of-mind centroid, since the intentional idiom of belief, desire, and reasons-versus-causes travels with it rather than being read off a physical loop the way feedback is. Transfer evidence is concrete (4): the belief-desire-intention analysis ports directly to reward-model-policy, the rational-agent model ports through bounded-rationality variants, and the adversarial-response prediction (agents game imposed metrics, objects do not) carries across tax enforcement, content moderation, and sanctions. The composite of 4 records a pattern recognized across nearly every domain, held just short of 5 by the human-and-animal-practice home that supplies half its working vocabulary.
- Composite substrate independence — 4 / 5
- Domain breadth — 5 / 5
- Structural abstraction — 4 / 5
- Transfer evidence — 4 / 5
Relationships to Other Abstractions¶
Current abstraction Agency Prime
Foundational — no parent edges in the catalog.
Children (15) — more specific cases that build on this
-
Internet Bot Domain-specific is a kind of Agency
Internet Bot strictly specializes
prime:agency.A qualifying bot has a delegated objective, task-relevant representation of remote state, and an action-selection coupling that chooses Internet operations expected to advance the objective. It occupies the software/network substrate and may have only minimal agency, but its response-sensitive loop is more than a passive mechanism. It is related todomain_specific:request_response, because many actions are exchanges with remote services, but a bot is an actor spanning many exchanges rather than the exchange pattern itself. It is related todomain_specific:access_endpoint, because endpoints are action surfaces, not because a route is a bot.prime:anthropomorphismand domain-specific social-bot detection become relevant when human-like cues shape attribution, yet human imitation is optional. Autonomy is not proposed as a parent. The operator ordinarily supplies the bot's goals and can stop or reconfigure it; execution without per-action control is operational independence, not necessarily self-government in the prime's stronger inner-authority sense. -
Language learning strategies Domain-specific is a kind of Agency
The proposed strict upward parent is
prime:agency.A learner literally selects actions in light of a represented language goal and beliefs about the task; linguistic targets, cognitive and metacognitive variants, contextual fit, measurement, and instruction supply the autonomous domain residual. The edge is proposal-only and points to a frozen prior-baseline Prime. The entry does not collapse into the parent because learner-selected goal-directed operations applied to additional-language learning or use under contextual and measurement boundaries, rather than a teaching method, fixed trait, generic study tip, automatic competence, or any behavior correlated with proficiency A thematic neighbor is declined whenever it does not literally subsume that rule. The prospective workspace queue contains one strict upward edge toprime:agency. No live DAG mutation is authorized. -
Libertarianism Domain-specific is a kind of Agency
The proposed strict upward parent is
prime:agency.prime:agency is the nearest broader Prime; the source-domain carrier and recognition invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Libertarianism adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity fixed by the thinker and tradition, definition of liberty and personhood, account of rights and nonaggression, property and economic theory, legitimate state or nonstate authority, equality and rectification, defense and enforcement and institutional implications are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Libertarianism. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge toprime:agency. No live DAG mutation is authorized.
- Prohairesis Domain-specific is a kind of Agency
The proposed strict upward parent is `prime:agency`.prime:agency is the nearest broader Prime; the source domain and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Prohairesis adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity determined by the philosopher, text, period and translation, impression or deliberative object, rational assessment, assent, desire and intention, voluntary-control boundary, action relation, moral evaluation, and differences between Aristotelian and Stoic use are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Prohairesis. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:agency`. No live DAG mutation is authorized.
- Purposive behaviorism Domain-specific is a kind of Agency
The proposed strict upward parent is `prime:agency`.prime:agency is the nearest broader Prime; the source domain and invariant supply the autonomous residual. This is a proposal-only workspace relationship: the accepted Prime supplies a genuinely instantiated structural prerequisite or superclass, while Purposive behaviorism adds domain-specific constraints. The entry does not collapse into that parent because the domain-specific identity determined by the historical theory and text, organism and environment, molar behavior, goal object, learned cue–cue or means–end expectancy, reinforcement condition and discriminating evidence are explicit It also declines a nearby thematic catalog node: the neighbor does not literally subsume the constitutive identity of Purposive behaviorism. This explicit assert-and-decline pattern keeps the proposed DAG narrow and prevents a merely thematic edge. The prospective workspace queue contains one strict upward edge to `prime:agency`. No live DAG mutation is authorized.
- Entrepreneurship Domain-specific presupposes Agency
**Agency** is the minimal prerequisite: entrepreneurial action requires an actor with a represented aim, beliefs about a situation, and action selection responsive to those beliefs.The proposed DAG relation therefore uses strict compositional presupposition rather than subsumption; entrepreneurship is not a kind of substrate-neutral agency. **Decision** appears at the pursuit threshold and at later continuation, pivot, and exit points. **Mobilization** describes activation and channeling of people, capital, knowledge, and relationships, but entrepreneurship need not begin with a latent reservoir or include demobilization. **Creative Destruction** describes one macroeconomic consequence of innovative entry, not every entrepreneurial episode. **Feedback**, **Learning**, and **Uncertainty** supply mechanisms without closing the domain identity. Narrower domain nodes include Entrepreneurial Discovery, Effectuation, and Entrepreneurial Bricolage. They should remain distinct rather than become aliases: one may discover without using effectual logic, use effectuation without founding a venture, or practice bricolage within entrepreneurial or non-entrepreneurial work.
- Existential Angst Domain-specific presupposes Agency
Existential angst presupposes an agent capable of representing possible lives and treating its own choices as reasons-responsive actions for which it bears responsibility.Agency supplies the prerequisite condition: A system pursues representable goals through actions whose selection is sensitive to its beliefs about its situation, via a goal-representation, world-model, and action-selection coupling. Existential Angst operates against that background: Anxiety from lack of meaning. If the parent condition is removed, the child relation becomes undefined or loses the mechanism asserted by this edge; the parent can obtain independently, so the relation is presupposition rather than subsumption.
- Female gaze Domain-specific presupposes Agency
The accepted reference-grade review places Female gaze under Agency because the child instantiates or depends on the parent's broader structure while retaining its own constitutive identity.An interpretive and creative framework representing women as agents and subjects through the spectator's, character's, or creator's gaze rather than as passive objects. The parent is defined more broadly: A system pursues representable goals through actions whose selection is sensitive to its beliefs about its situation, via a goal-representation, world-model, and action-selection coupling.
- Information Avoidance Domain-specific presupposes Agency
Information avoidance presupposes agency because it requires a system to forecast the consequence of knowing, pursue a preserved benefit, and select an action that prevents signal receipt.The source makes the agentive boundary explicit: without a utility-bearing anticipator there is no one for whom knowing is worse. The avoider has a represented goal such as affect relief or obligation escape, a model of what the signal is likely to reveal, and an action-selection coupling that chooses not opening, not asking, delegating, or designing the signal away.
- Information Seeking Domain-specific presupposes Agency
The accepted reference-grade review places Information Seeking under Agency because the child instantiates or depends on the parent's broader structure while retaining its own constitutive identity.Goal-directed human activity that begins with a perceived information need and proceeds through source selection, searching, encountering, evaluating, using, and sometimes abandoning information across formal and informal channels. The parent is defined more broadly: A system pursues representable goals through actions whose selection is sensitive to its beliefs about its situation, via a goal-representation, world-model, and action-selection coupling.
- Uses and Gratifications Domain-specific is part of Agency
An active audience member with goals and a model of available media is a strict constituent of Uses and Gratifications.The framework's defining inversion rejects a passive recipient and requires an actor whose anticipated outcomes guide content choice. Without that goal-sensitive action selection, the theory collapses back into an effects-only account.
- Agency Problem Prime presupposes Agency
Agency_problem (principal-agent) PRESUPPOSES two agents, each satisfying agency's goal/world-model/action-selection triad.agency_problem (principal-agent) PRESUPPOSES two agents, each satisfying agency's goal/world-model/action-selection triad. It is a relational configuration built ON agency, not a kind-of agency — so the edge flavor is presupposes, not is-a.
- Anthropomorphism Prime presupposes Agency
Anthropomorphism presupposes a concept of agency because it projects an agentive interpretation onto a target whose agency has not been established.Without an observer-side category of goal-directed agency, surface cues could prompt resemblance or pattern recognition but could not support the distinctive inference that the target intends, understands, wants, or acts for reasons.
- Co-Presence Prime presupposes Agency
Co-presence presupposes agents capable of detecting one another and selecting responses to that mutually available state.Mere simultaneous existence is not co-presence. At least two bounded systems must be able to register the other's availability and condition possible response on that registration. Agency supplies the responsive entities whose mutual availability constitutes the relation.
- Self-Efficacy Prime presupposes Agency
Self_efficacy is an agent's second-order belief about its OWN capacity to act — presupposes the first-order agency triad.Agency supplies the prerequisite condition: A system pursues representable goals through actions whose selection is sensitive to its beliefs about its situation, via a goal-representation, world-model, and action-selection coupling. Self-Efficacy operates against that background: Belief in capability. If the parent condition is removed, the child relation becomes undefined or loses the mechanism asserted by this edge; the parent can obtain independently, so the relation is presupposition rather than subsumption.
Neighborhood in Abstraction Space¶
Agency sits among the more crowded primes in the catalog (31st percentile for distinctiveness): several abstractions describe nearly the same structure, so a description that fits it will tend to fit its neighbors too — transporting it usually means disambiguating within this family rather than landing on it exactly.
Family — Unclustered & Miscellaneous (424 primes)
Nearest neighbors
- Good Regulator Theorem — 0.74
- Mission Command — 0.73
- Shared Mental Model — 0.72
- Mental Model — 0.72
- Self-Defeating Prediction — 0.72
Computed from structural-signature embeddings · 2026-09-10
Not to Be Confused With¶
The most pressing confusion — and the one the near-identical name forces — is with agency_problem (similarity 0.999). They are not the same concept at different resolutions; they sit at different levels of the dependency stack. Agency is the substrate: the property of being a goal-pursuing, model-using, action-selecting system at all. The agency problem (the principal-agent problem) is a configuration built on top of agency: it presupposes two agents — a principal and a delegate — and studies what happens when the delegate's goal-representation diverges from the principal's while the principal cannot fully observe the delegate's actions. Every component of the agency problem — delegation, hidden action, incentive misalignment, monitoring cost — assumes agency already established in both parties; none of it makes sense applied to objects. The clean claim is parent/child: agency is necessary for an agency problem, but agency without delegation, without a second party, and without goal divergence is simply an agent acting, with no "problem" at all. A solitary forager, a single RL agent in a gridworld, has full agency and no agency problem. Collapsing the two would erase the distinction between "is this a goal-pursuing system?" and "are two goal-pursuing systems misaligned under hidden action?" — questions with entirely different diagnostics and interventions.
A second genuine confusion is with autonomy. Both attach to goal-pursuing systems and both carry normative weight, but they answer different questions. Agency asks whether the system pursues represented goals through a model-mediated coupling — the architecture. Autonomy asks whose goals those are: whether the system governs itself by ends it has set or endorsed, versus executing ends imposed from outside. The two dissociate cleanly. A drone executing a commander's targeting goal has agency (goal, model, action-selection, re-planning) but little autonomy (the goal is wholly imposed). A person under coercion retains agency while their autonomy is compromised. Conversely, granting a system more autonomy does not add agency components; it relocates the source of the goal-representation from outside to inside. Confusing them leads to the error of thinking that constraining a system's autonomy removes its agency — and so treating a constrained-but-agentic system as an object, which the prime warns invites adversarial surprise.
A third confusion is with teleology. Teleology is the attribution of goal-directedness as an explanatory gloss, applicable to systems with no internal goal-representation whatsoever: the heart is "for" pumping blood, evolution "designs" the eye, a river "seeks" the sea. Agency requires more — the goal must be represented inside the system and must drive action selection through a world-model. The distinction is exactly the gap between a functional/as-if story and a mechanistic claim about internal architecture. Many over-attributions of agency are really teleological glosses promoted to literal agency: reading purpose into a market, a gene, or a weather system. The discriminating test is whether removing the alleged goal-representation would change the system's behavior (agency) or only the observer's description of it (teleology).
For a practitioner the stakes of these distinctions are operational. Confusing agency with the agency problem conflates "is this an agent?" with "are these agents misaligned?" — skipping the prior attribution question that determines whether force or reasons is the right tool. Confusing agency with autonomy leads to treating constrained agents as controllable objects, inviting the gamed-metric surprise. Confusing agency with teleology promotes explanatory glosses into intervention targets, wasting effort persuading systems that cannot represent goals. The unifying discipline is to settle, in order, three questions: does the system have the triad (agency), are its goals its own (autonomy), and is the goal really represented inside it or only ascribed by me (teleology versus agency) — because each answer changes whether, and how, the system can be influenced.
Solution Archetypes¶
Solution archetypes in the catalog that build on this prime — directly (this prime is a source ingredient) or as a related prime.
Built directly on this prime (8)
- Affordance Shaping: Arrange the fit between an agent and its environment so the right actions are available, noticeable, and easier at the moment they matter.▸ Mechanisms (12)
- Affordance Audit — Systematically inventories what an existing environment actually affords, to whom, and where its usable actions diverge from the intended ones.
- Contextual Inquiry or Walkthrough — Learns what agents are really trying to do — and where they go wrong — by observing them in their own context of use rather than reasoning about them from a desk.
- Desire Path Observation — Reads the tracks agents wear into an environment as evidence of the route they actually want, revealing the affordance the design should have offered.
- Friction Adjustment — Tilts the effort, steps, and salience of a path — smoothing the desired action and adding drag to the harmful one — so behaviour changes without persuasion or prohibition.
- Physical or Digital Keying — Shapes the environment so that only the correct action physically or logically fits, making the harmful move impossible rather than merely discouraged.
- Prototype A/B or Multivariate Test — Puts two or more candidate shapings in front of real agents at once and lets their measured behaviour decide which one actually moves the target action.
- Robot Action-Space Mapping — Maps the actions a robot can actually execute in its environment — reachable, collision-free, within its own limits — so the intended action lies inside the feasible space and the harmful ones fall outside it.
- Safe Default or Preselected Path — Makes the desired, low-risk option the one that happens when the agent does nothing — while keeping the alternative one easy, reversible step away.
- Signifier Prototyping — Designs and iterates the perceivable cues that tell an agent an action is available and how to perform it — turning a hidden affordance into an obvious one for everyone who has to see it.
- Task and Capability Analysis — Decomposes the goal into the actions it requires and checks each against what the agent can actually perceive, reach, and do — locating where the task outruns the agent's capability.
- Usability or Field Test — Puts the shaped affordance in front of real users in a realistic setting and records what they actually do, so the design is judged by behaviour rather than by the designer's intention.
- Wayfinding Marker — Places perceivable signals along a route so the correct next step is always the salient one, revealing the path a piece at a time instead of demanding a whole map.
- Agency / Structure Attribution Balance: Attribute outcomes to actors and structures through explicit causal roles, opportunity conditions, cross-scale evidence, and counterfactual tests.▸ Mechanisms (10)
- Actor–Structure Evidence Matrix — Arrays every actor action against every structural condition in one grid, forcing each cell to carry dated evidence, a causal role, and an effect so that unsupported claims and empty cells become visible.
- Comparative-Case Attribution Test — Uses real cases where the actor varied under similar structures, or the structure varied under similar actors, as natural experiments that discriminate among competing attributions.
- Counterfactual Actor-Substitution Probe — Holds the structure fixed and asks what a plausible substitute, an absence, or a delay of the focal actor would have changed — bounding how replaceable the person really was.
- Narrative Focal-Actor Audit — Compares how much of the record an actor occupies against how much causal evidence actually supports them, surfacing archive and commemoration bias and the contributors the record erased.
- Network and Institutional Position Map — Charts positions, authority, brokerage, and selection in the relational structure to show which options a role afforded and how replaceable its occupant was.
- Paired Micro/Macro Timeline — Runs a micro timeline of dated actor events beside a macro timeline of slow structural change so that contingency and pattern can be read against each other and an actor's timing role becomes visible.
- Plural Causal Synthesis Review — Holds several competing causal models open at once, restates the scoped target, and publishes a bounded, versioned synthesis that separates causal contribution from praise and blame.
- Process Tracing Across Levels — Reconstructs the causal chain by which a single decision propagates upward through individual, group, organizational, and institutional levels to the macro outcome, testing each link against the evidence it would have to leave.
- Structural-Constraint Relaxation Probe — Holds the actor fixed and loosens or tightens one structural condition at a time to see whether the actor's effect survives the change — measuring how conditional that effect really was.
- Turning-Point Opportunity Analysis — Tests whether a narrow option window or an inherited path made one decision unusually consequential, and how quickly the window opened and closed.
- Agentic Control Loop Design: Agency becomes real when goals, situation models, available actions, authority, execution, feedback, and learning are coupled into a loop that can intentionally change outcomes.▸ Mechanisms (10)
- Action-Effect Feedback Review — A recurring review that attributes what an action did and did not change, updating the actor's read on what is now within their control.
- After-Action Learning Cycle — A recurring, blame-free review that turns what actually happened into concrete revisions of the model and the next action.
- Agency Health Dashboard — Turns the live health of an agency loop — is feedback timely, is the actor actually acting, is discretion being used — into a small set of continuously-watched signals.
- Agency Loop Map — Lays the agent's full goal-to-feedback loop out as one connected diagram so a missing or broken coupling becomes visible at a glance.
- Briefback or Intent Confirmation — Before acting, the actor restates the goal, constraints, and plan back to the tasker to confirm shared understanding and surface conflicts early.
- Controllability Mapping Checklist — Sorts a situation into controllable, influenceable, constrained, and uncontrollable parts before any action is chosen.
- Decision-Rights Matrix — Maps each class of decision to who may decide, approve, be consulted, or merely be informed — fixing the agent's authority before any single choice arises.
- Graduated Autonomy Ramp — A staged schedule that widens an actor's decision authority as evidence of competence accumulates, with support fading as autonomy grows.
- Model Assumption Register — A living list of every assumption the agent's world model rests on, each with an owner, a confidence, and a stated trigger for when it must be revisited.
- Safe Action Menu — A fixed template of pre-approved, in-bounds actions for a high-risk setting, with an escalate path for anything the menu does not cover.
- Agent–Environment Co-Shaping: Shape the environment an agent or population inhabits so the resulting conditions improve future behavior and adaptation—and keep governing the feedback as both sides change.▸ Mechanisms (12)
- Adaptive Management Cycle — Governs a co-shaping environment as a running act→monitor→learn→adjust loop, updating the intervention from evidence as agents and their surroundings keep changing each other.
- Agent-Based Niche Simulation — Runs the co-shaping loop forward in silico with many adaptive agents, so you can watch which environmental changes stay viable — and which get gamed — before committing them for real.
- Causal-Loop and Environment-State Map — A single diagram of the environment's boundary, its state variables, and the reinforcing and balancing feedback loops through which agents and their surroundings change each other.
- Ecological Restoration Pilot — A bounded field intervention that jump-starts a self-sustaining successional trajectory in a degraded habitat, then hands the recovery over to the system's own feedbacks.
- Environmental Indicator Dashboard — A live instrument panel that tracks how agents and their environment are co-adapting — and flags when someone is adapting to game the very signals you steer by.
- Habitat or Spatial Reconfiguration — Rearranges physical space so its new adjacencies, sightlines, and barriers quietly reshape how the people or organisms moving through it behave.
- Infrastructure and Default Redesign — Rebuilds the shared substrate and default settings people act within, so the behaviour you want becomes the path of least resistance instead of an act of willpower.
- Institutional Rule and Incentive Redesign — Rewrites the rules, sanctions, and payoffs of a shared setting so the environment itself selects for the behaviour you want — and those who act bear its consequences.
- Legacy and Maintenance Register — Keeps a standing record of what past shaping left behind — the constructions, dependencies, and obligations later agents inherit — so nothing load-bearing is forgotten, retired blindly, or left to rot.
- Platform-Ecosystem Rule Change — Changes the rules of a live digital ecosystem and governs the fast, often adversarial way participants re-adapt to them.
- Staged Reversible Environment Pilot — Tests an environmental change on a bounded, undoable slice first — keeping an escape path and preserving options — so you learn what it does before it hardens into something you can't take back.
- Stakeholder Boundary Review — Decides who counts as inside the system being shaped — constructors, beneficiaries, and the affected outsiders who bear the spillovers — before the boundary is drawn implicitly by whoever holds the pen.
- Alienation Reconnection: Reconnect people to agency, meaning, community, contribution, and system consequences when structures make participation feel remote, opaque, or powerless.▸ Mechanisms (10)
- Alienation Relation-Mapping Workshop — Convenes the people a system estranges to name, together, exactly which links — to agency, meaning, contribution, feedback, belonging, governance — the structure has severed, turning a diffuse sense of powerlessness into a mapped set of nameable broken relations.
- Bounded Decision-Rights Charter — Hands participants a small but genuine zone of decisions they own outright — binding, not advisory — with the right to refuse or exit intact, converting nominal 'input' into consequential agency.
- Closed-Loop Response Commitment — Guarantees that every signal a participant sends gets a tracked response within a bounded time — acted on, or explicitly declined with a reason — and that the sender is told what changed and credited for it, so feedback stops disappearing into a void.
- Contribution-to-Beneficiary Review — Traces a participant's work to the actual person it reaches and stages direct contact between them, restoring the last-mile line of sight from effort to a human benefit that the system usually hides.
- Mediation-Layer Transparency Review — Exposes the stack of intermediaries — algorithms, gatekeepers, layers of process — sitting between a participant and the system, judges which add value versus only distance, and plans to make the necessary ones legible and cut the rest.
- Participant Governance Forum — A standing body in which affected participants hold seats and real decision power over matters that concern them, with formal routes to contest and reshape policy — giving collective voice a durable venue with teeth, not a suggestion box.
- Participant Journey and Consequence Trace — Follows one participant's action all the way through the system to its real consequence and back again, built from their own account, so the exact segment where the line of sight goes dark becomes visible and locatable.
- Peer and Steward Connection Circle — A small, recurring circle of peers anchored by a designated steward, giving isolated participants a durable human place to belong and a guide through the disorientation of reconnection or change.
- Reconnection Pulse and Burden Audit — Periodically measures whether reconnection is actually taking hold — and whether the effort to reconnect has itself become a burden — catching drift and participation fatigue before they quietly undo the gains.
- Whole-System Context Session — Lays out the whole system a participant's work sits inside — its purpose, its parts, and where they fit — so a fragment of a job regains its meaning within the larger whole and a shared purpose knits contributors into a collective.
- Other-Agent State Model Calibration: Model another agent as having its own partial knowledge, goals, attention, constraints, and interpretations, then update that model from evidence before routing action through it.▸ Mechanisms (11)
- Active Listening Loop — Reflects the other agent's meaning back to them and invites correction, so the actor's model is checked and repaired live — in the exchange — rather than after the misunderstanding lands.
- Belief-Desire-Knowledge Map — Lays out what another agent probably believes, wants, knows, lacks, fears, and expects as an explicit set of hypotheses, each carrying a confidence level.
- Consent and Privacy Boundary Checklist — Gates whether it is legitimate to build, keep, share, and act on a model of another agent's private state — before the model is used, not after.
- Counterparty Model Red Team — Attacks a working model of a strategic counterparty by manufacturing rival explanations for their motives, constraints, and moves, to break the single story the actor has settled on.
- Empathy Map with Evidence Marks — Captures what another agent seems to see, hear, think, feel, say, and do — with every cell tagged as observed evidence or actor assumption.
- False-Belief Check — Tests the single assumption that the other agent knows what you know — catching curse-of-knowledge errors before they distort an explanation, interface, or instruction.
- Interaction After-Action Review — A recurring retrospective that asks where the model of the other agent helped, failed, surprised, or harmed — and rewrites the interaction rules accordingly.
- Perspective-Taking Interview — Replaces inference with direct, open-ended questioning to learn the other agent's actual understanding, constraints, and priorities.
- Prediction and Surprise Log — A running record of what the other agent was predicted to do, what they actually did, and how the model changed — making calibration visible across repeated interactions.
- Role-Reversal Simulation — Steps through the situation from the other agent's information, constraints, and incentives — arguing their case as they would — to expose where the actor's model is really just projection.
- Stakeholder Hidden-Constraint Board — A shared visual board that names each stakeholder and makes their invisible constraints, fears, incentives, and information gaps explicit for a team to design around.
- Role Expectation Architecture: When coordination depends on a recurring social position, design the role as a clear, occupiable bundle of expected behaviours, authority, obligations, interfaces, support, conflict guards, and handoff rules.▸ Mechanisms (12)
- Conflict-of-Interest Disclosure — Makes a decision-maker declare the relationships and incentives that could skew their judgment, so a specific decision can be checked for independence.
- Delegation Letter or Authority Envelope — Transfers a bounded, revocable slice of decision authority to a named holder — stating exactly what they may decide, up to what limit, and what to do at the edge of that envelope.
- Handoff Checklist — A structured transfer list that moves a role from an outgoing holder to a successor without dropping open commitments, live context, or hard-won know-how.
- Onboarding and Role Shadowing Runbook — A structured ramp that brings a new holder up to a role's competence bar by provisioning support and mentorship and by having them learn through supervised shadowing of an experienced holder.
- Position Description or Office Mandate — The founding document that establishes a position exists, states what its holder is responsible for and owes to others, and makes the role recognizable independent of whoever currently fills it.
- RACI or Decision Participation Matrix — Lays every recurring task or decision against every role in a grid and tags each cell, so exactly one role is Accountable and no decision right is left blank or doubled.
- Role Card or Participation Card — A single-role, at-a-glance card — this position, the few things you do, the near ones you don't, and whom you serve — small enough to hand someone the moment they step into the seat.
- Role Charter — Constitutes a role or governing body as a legitimate office — fixing its remit and decision authority, the path by which it answers for its actions, and how it is properly filled and vacated.
- Role Compatibility Check — A pre-appointment screen that tests a proposed role assignment against the role's competence bar and against conflict and separation constraints, before the assignment is made.
- Role Review Retrospective — A recurring session that puts the role itself — not the person in it — on the table: is it still needed, still sane in scope, still bearable, and what should change?
- Role Rotation or Deputy Schedule — A standing schedule of who holds a role now, who covers when they're out, and who takes over next — so the position survives any single person leaving the seat.
- Swimlane or Service Blueprint — Draws the work as parallel lanes — one per role — so every step, handoff, and 'whose job is this?' gap shows up as a line crossing (or failing to cross) a lane boundary.
- Tool-Repertoire Bias Counterbalancing: Counter tool-induced problem bias by describing the need before choosing the tool, mapping what the tool can and cannot grip, testing alternative instruments, and creating a path for residual cases.▸ Mechanisms (9)
- Affordance Blind-Spot Walkthrough — Walks a single tool through what it makes easy, hard, impossible, visible, and invisible — surfacing the residuals its grip would otherwise erase.
- Alternative-Tool Red Team — Deliberately re-describes the problem from a rival discipline, representation, or stakeholder position to prove whether the default tool is genuinely fit or merely familiar.
- Borrow-or-Refer Protocol — Routes a case to borrowed capability, referral, or escalation when the local repertoire cannot fit it, so tool limits never become the boundary of responsibility.
- Favored-Tool Pause Rule — A circuit-breaker that blocks reflexive use of the favored method until its fit to the restated problem has been established.
- Problem-First Intake Template — Captures the need — outcome, stakeholders, constraints, evidence, uncertainty — in problem-first language and fixes how success will be judged, all before any tool, method, or category is named.
- Representation Fit Scorecard — Scores how well each candidate representation preserves the problem-first need across explicit fit dimensions, making tool choices comparable instead of habitual.
- Residual Case Log — A standing register of cases that don't fit current categories, each tagged with an owner and a route, so residuals are preserved and resolved rather than dumped.
- Tool Repertoire Inventory — Makes the local toolset visible as a maintained catalog of methods, systems, credentials, and habits, and keeps it current so the repertoire can be seen as a bias source rather than the whole world.
- Tool-Mismatch Postmortem — After a tool-native success masks a real-world failure, reconstructs the mismatch and feeds it back into training, procurement, and repertoire expansion.
Also a related prime in 15 archetypes
- Advantageous Repositioning: Gain advantage by moving to a better position in the option, terrain, timing, information, or institutional space instead of fighting the same contest from a worse position.
- Coercive Leverage Governance: Use explicit, bounded consequences to reshape another actor's choice set while preserving legitimacy, proportionality, verification, and an exit from coercive pressure when conditions are met.
- Conformity Pressure Calibration: Calibrate the pressure to match a group standard by protecting private judgment, exposing social-pressure channels, and preserving safe divergence before alignment becomes automatic.
- Holonic Autonomy Nesting: Design nested units as autonomous local wholes and dependent parts at the same time, with explicit boundaries, interfaces, escalation paths, and cross-level invariants.
- Human-Capacity Accommodation Design: Diagnose the mismatch between human capacity and system demand, then change the task, environment, interface, timing, modality, or support so people can achieve essential outcomes safely and with dignity.
- Model-Based Regulation: Embed a decision-relevant, continuously tested model of the system inside its regulator so interventions are state-aware, predictive, auditable, and revisable.
- Narrative Transportation Persuasion Design: Use a storyworld to let the audience experience the target belief or attitude as lived consequence rather than as a bare proposition, then make the resulting shift ethically inspectable.
- Outcome Responsibility Attribution Calibration: Assign credit or blame only after separating outcome, causal contribution, control, duty, knowledge, and uncertainty.
- Perception-Comprehension-Projection Loop Design: Keep action aligned with a moving situation by continuously refreshing what is seen, what it means, what is likely next, and what decision it now supports.
- Principal-Bound Authority Mediation: Let a deputy act only when the requesting principal, stated intent, delegated scope, and use of the deputy’s authority are explicitly bound and checkable.
References¶
[1] Bratman, Michael E. Intention, Plans, and Practical Reason. Cambridge, MA: Harvard University Press, 1987. Canonical belief-desire-intention (BDI) analysis of intentional action and the planning theory of agency. registry ↩
[2] Russell, Stuart. Human Compatible: Artificial Intelligence and the Problem of Control. New York: Viking, 2019. Frames AI agents in terms of objectives, world-models, and action selection and the alignment problem; supports reward-gaming, deception, and emergent sub-goal debates as one framework. (See also Sutton & Barto for the reward-model-policy formalism.) registry ↩a ↩b
[3] Simon, Herbert A. "A Behavioral Model of Rational Choice." Quarterly Journal of Economics, vol. 69, no. 1 (1955): 99–118. Introduces bounded rationality — real agents approximating the rational-agent model under resource limits. registry ↩
[4] Hart, H. L. A. Punishment and Responsibility: Essays in the Philosophy of Law. Oxford: Oxford University Press, 1968. Canonical legal-philosophy treatment of criminal responsibility, excusing conditions, and the capacity conditions (volition, knowledge, control) on which ascriptions of responsibility — and thus the agent/patient distinction — depend. registry ↩
[5] Giddens, Anthony. The Constitution of Society: Outline of the Theory of Structuration. Cambridge: Polity Press, 1984. Canonical statement of the structure-versus-agency debate in social theory. registry ↩
[6] Beauchamp, Tom L., and James F. Childress. Principles of Biomedical Ethics, 8th ed. New York: Oxford University Press, 2019. Operationalizes capacity to consent as informed, voluntary, and competent — the components mapping onto belief, goal, and action-selection. registry ↩
[7] Sutton, Richard S., and Andrew G. Barto. Reinforcement Learning: An Introduction, 2nd ed. Cambridge, MA: MIT Press, 2018. Standard formalization of the reward function R(s,a), transition model, policy π, expected discounted return, value iteration, and discount factor γ — the RL instantiation of the goal/world-model/action-selection triad. registry ↩a ↩b ↩c ↩d
[8] Manheim, David, and Scott Garrabrant. "Categorizing Variants of Goodhart's Law." arXiv:1803.04585, 2018. Supports the reward-hacking / gamed-metric prediction that agents re-plan adversarially against imposed measures. (See also Amodei et al., "Concrete Problems in AI Safety," arXiv:1606.06565, 2016, on reward hacking.) registry ↩