Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Reinforcement learning with value advice

Loading...
Thumbnail Image

Authors

Daswani, Mayank
Sunehag, Peter
Hutter, Marcus

Journal Title

Journal ISSN

Volume Title

Publisher

Journal of Machine Learning Research

Abstract

The problem we consider in this paper is reinforcement learning with value advice. In this setting, the agent is given limited access to an oracle that can tell it the expected return (value) of any state-action pair with respect to the optimal policy. The agent must use this value to learn an explicit policy that performs well in the environment. We provide an algorithm called RLAdvice, based on the imitation learning algorithm DAgger. We illustrate the effectiveness of this method in the Arcade Learning Environment on three different games, using value estimates from UCT as advice.

Description

Citation

Source

Book Title

Proceedings of the 6th Asian Conference on Machine Learning

Entity type

Access Statement

Open Access

License Rights

Creative Commons Attribution licence

DOI

Restricted until