Markov Decision Processes¶
Puterman, M. L. (1994). Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley.
Cited by¶
2 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Markov Decision Processes (MDPs)
This sourceCanonical operations-research reference on MDPs; comprehensive treatment of finite/infinite-horizon, discounted/average-reward, and structured-policy results, with applications across inventory, maintenance, and queueing.
- Optimal Stopping Rule
- Not the full sequential-control apparatus of `markov_decision_processes_mdps`. An MDP optimizes a policy over actions and rewards across many states; optimal stopping is the special case where the only action is "stop or continue."
This sourceCanonical treatment of MDPs as policy optimization over actions and rewards across states; optimal stopping is the degenerate case with a stop/continue action set.
- Not the full sequential-control apparatus of `markov_decision_processes_mdps`. An MDP optimizes a policy over actions and rewards across many states; optimal stopping is the special case where the only action is "stop or continue."
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:430d6b73b15d · see in the full table