Doubly Robust Policy Evaluation and Learning¶
Dudík, M., Langford, & Li, L. (2011). Doubly Robust Policy Evaluation and Learning. Proceedings of the 28th International Conference on Machine Learning, 1097-1104.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Mechanisms¶
- Historical Replay
- Its deep limit is that logged history only records the outcomes of the actions that were actually taken, so replay cannot truly observe how the world would have responded to the new policy's different choices — the missing-counterfactual problem at the heart of off-policy evaluation.
This sourceFrames off-policy evaluation as estimating a new policy from historical contexts, chosen actions, and only the rewards those actions revealed, leaving outcomes under different actions unobserved.
- Its deep limit is that logged history only records the outcomes of the actions that were actually taken, so replay cannot truly observe how the world would have responded to the new policy's different choices — the missing-counterfactual problem at the heart of off-policy evaluation.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:1775f6c2aa66 · see in the full table