• Classified by Topic • Classified by Publication Type • Sorted by Date • Sorted by First Author Last Name • Classified by Funding Source •
Framing reinforcement learning from human reward: Reward positivity, temporal discounting, episodicity, and performance.
W. Bradley Knox and Peter Stone.
Artificial
Intelligence, 225(), August 2015.
Artificial
Intelligence
Contains material that was previously published in a IUI 2013 paper and a RoMan 2012 paper that was nominated as a CoTeSys Cognitive Robotics BEST PAPER AWARD FINALIST.
Several studies have demonstrated that reward from a human trainer can be a powerful feedback signal for control-learning algorithms. However, the space of algorithms for learning from such human reward has hitherto not been explored systematically. Using model-based reinforcement learning from human reward, this article investigates the problem of learning from human reward through six experiments, focusing on the relationships between reward positivity, which is how generally positive a trainer's reward values are; temporal discounting, the extent to which future reward is discounted in value; episodicity, whether task learning occurs in discrete learning episodes instead of one continuing session; and task performance, the agent's performance on the task the trainer intends to teach. This investigation is motivated by the observation that an agent can pursue different learning objectives, leading to different resulting behaviors. We search for learning objectives that lead the agent to behave as the trainer intends.We identify and empirically support a ``positive circuits’’ problem with low discounting (i.e., high discount factors) for episodic, goal-based tasks that arises from an observed bias among humans towards giving positive reward, resulting in an endorsement of myopic learning for such domains. We then show that converting simple episodic tasks to be non-episodic (i.e., continuing) reduces and in some cases resolves issues present in episodic tasks with generally positive reward and—relatedly—enables highly successful learning with non-myopic valuation in multiple user studies. The primary learning algorithm introduced in this article, which we call ``vi-tame’’, is the first algorithm to successfully learn non-myopically from reward generated by a human trainer; we also empirically show that such non-myopic valuation facilitates higher-level understanding of the task. Anticipating the complexity of real-world problems, we perform further studies—one with a failure state added—that compare (1) learning when states are updated asynchronously with local bias—i.e., states quickly reachable from the agent's current state are updated more often than other states—to (2) learning with the fully synchronous sweeps across each state in the vi-tamer algorithm. With these locally biased updates, we find that the general positivity of human reward creates problems even for continuing tasks, revealing a distinct research challenge for future work.
@ARTICLE{AIJ15-Knox, AUTHOR={W.~Bradley Knox and Peter Stone}, TITLE={Framing reinforcement learning from human reward: Reward positivity, temporal discounting, episodicity, and performance}, JOURNAL={Artificial Intelligence}, VOLUME={225}, YEAR={2015}, month={August}, NUMBER={}, URL={http://www.sciencedirect.com/science/article/pii/S0004370215000557}, DOI={10.1016/j.artint.2015.03.009}, ISSN={}, ABSTRACT={Several studies have demonstrated that reward from a human trainer can be a powerful feedback signal for control-learning algorithms. However, the space of algorithms for learning from such human reward has hitherto not been explored systematically. Using model-based reinforcement learning from human reward, this article investigates the problem of learning from human reward through six experiments, focusing on the relationships between reward positivity, which is how generally positive a trainer's reward values are; temporal discounting, the extent to which future reward is discounted in value; episodicity, whether task learning occurs in discrete learning episodes instead of one continuing session; and task performance, the agent's performance on the task the trainer intends to teach. This investigation is motivated by the observation that an agent can pursue different learning objectives, leading to different resulting behaviors. We search for learning objectives that lead the agent to behave as the trainer intends. We identify and empirically support a ``positive circuitsââ problem with low discounting (i.e., high discount factors) for episodic, goal-based tasks that arises from an observed bias among humans towards giving positive reward, resulting in an endorsement of myopic learning for such domains. We then show that converting simple episodic tasks to be non-episodic (i.e., continuing) reduces and in some cases resolves issues present in episodic tasks with generally positive reward andârelatedlyâenables highly successful learning with non-myopic valuation in multiple user studies. The primary learning algorithm introduced in this article, which we call ``vi-tameââ, is the first algorithm to successfully learn non-myopically from reward generated by a human trainer; we also empirically show that such non-myopic valuation facilitates higher-level understanding of the task. Anticipating the complexity of real-world problems, we perform further studiesâone with a failure state addedâthat compare (1) learning when states are updated asynchronously with local biasâi.e., states quickly reachable from the agent's current state are updated more often than other statesâto (2) learning with the fully synchronous sweeps across each state in the vi-tamerâ algorithm. With these locally biased updates, we find that the general positivity of human reward creates problems even for continuing tasks, revealing a distinct research challenge for future work.}, wwwnote={<a href="http://www.journals.elsevier.com/artificial-intelligence/">Artificial Intelligence</a><br> }
Generated by bib2html.pl (written by Patrick Riley ) on Tue Nov 19, 2024 10:24:39