Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Feature Markov Decision Processes

dc.contributor.authorHutter, Marcus
dc.date.accessioned2015-08-26T05:33:54Z
dc.date.available2015-08-26T05:33:54Z
dc.date.issued2009-05
dc.description.abstractGeneral purpose intelligent learning agents cycle through (complex,non-MDP) sequences of observations, actions, and rewards. On the other hand, reinforcement learning is welldeveloped for small finite state Markov Decision Processes (MDPs). So far it is an art performed by human designers to extract the right state representation out of the bare observations, i.e. to reduce the agent setup to the MDP framework. Before we can think of mechanizing this search for suitable MDPs, we need a formal objective criterion. The main contribution of this article is to develop such a criterion. I also integrate the various parts into one learning algorithm. Extensions to more realistic dynamic Bayesian networks are developed in the companion article [Hut09].en_AU
dc.identifier.isbn9789078677246en_AU
dc.identifier.urihttp://hdl.handle.net/1885/14962
dc.publisherAtlantis Pressen_AU
dc.relation.ispartofArtificial general intelligence: proceedings of the second conference on Artificial General Intelligence, AGI 2009, Arlington, Virginia, USA, March 6-9, 2009en_AU
dc.rights© Atlantis Press. This article is distributed under the terms of the Creative Commons Attribution License, which permits non-commercial use, distribution and reproduction in any medium, provided the original work is properly cited.en_AU
dc.subjectReinforcement learningen_AU
dc.subjectMarkov decision processen_AU
dc.subjectpartial observabilityen_AU
dc.subjectfeature learningen_AU
dc.subjectexplore-exploiten_AU
dc.titleFeature Markov Decision Processesen_AU
dc.typeJournal articleen_AU
local.bibliographicCitation.lastpage6en_AU
local.bibliographicCitation.startpage1en_AU
local.contributor.affiliationHutter, M., Research School of Computer Science, The Australian National Universityen_AU
local.contributor.authoruidu4350841en_AU
local.identifier.doi10.2991/agi.2009.30en_AU
local.type.statusPublished Versionen_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Hutter Feature Markov Decision Processes 2009.pdf
Size:
420.46 KB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
884 B
Format:
Item-specific license agreed upon to submission
Description: