Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Exact Reduction of Huge Action Spaces in General Reinforcement Learning

dc.contributor.authorMajeed, Sultan
dc.contributor.authorHutter, Marcus
dc.coverage.spatialvirtual
dc.date.accessioned2024-01-29T23:01:34Z
dc.date.createdFebruary 2–9, 2021
dc.date.issued2021
dc.date.updated2022-10-02T07:18:42Z
dc.description.abstractThe reinforcement learning (RL) framework formalizes the notion of learning with interactions. Many real-world problems have large state-spaces and/or action-spaces such as in Go, StarCraft, protein folding, and robotics or are non-Markovian, which cause significant challenges to RL algorithms. In this work we address the large action-space problem by sequentializing actions, which can reduce the action-space size significantly, even down to two actions at the expense of an increased planning horizon. We provide explicit and exact constructions and equivalence proofs for all quantities of interest for arbitrary history-based processes. In the case of MDPs, this could help RL algorithms that bootstrap. In this work we show how action-binarization in the nonMDP case can significantly improve Extreme State Aggregation (ESA) bounds. ESA allows casting any (non-MDP, non-ergodic, history-based) RL problem into a fixed-sized non-Markovian state-space with the help of a surrogate Markovian process. On the upside, ESA enjoys similar optimality guarantees as Markovian models do. But a downside is that the size of the aggregated state-space becomes exponential in the size of the action-space. In this work, we patch this issue by binarizing the action-space. We provide an upper bound on the number of states of this binarized ESA that is logarithmic in the original action-space size, a double-exponential improvement.en_AU
dc.description.sponsorshipThis work has been supported by Australian Research Council grant DP150104590en_AU
dc.format.mimetypeapplication/pdfen_AU
dc.identifier.isbn978-1-57735-866-4en_AU
dc.identifier.urihttp://hdl.handle.net/1885/312404
dc.language.isoen_AUen_AU
dc.publisherThe AAAI Pressen_AU
dc.relationhttp://purl.org/au-research/grants/arc/DP150104590en_AU
dc.relation.ispartofseriesTHIRTY-FIFTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCEen_AU
dc.rights© 2021 Association for the Advancement of Artificial Intelligenceen_AU
dc.source.urihttps://ojs.aaai.org/index.php/AAAI/article/view/17074/16881en_AU
dc.titleExact Reduction of Huge Action Spaces in General Reinforcement Learningen_AU
dc.typeConference paperen_AU
dcterms.accessRightsFree Access via publisher websiteen_AU
local.bibliographicCitation.lastpage8883en_AU
local.bibliographicCitation.startpage8874en_AU
local.contributor.affiliationMajeed, Sultan, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationHutter, Marcus, College of Engineering and Computer Science, ANUen_AU
local.contributor.authoruidMajeed, Sultan, u5447242en_AU
local.contributor.authoruidHutter, Marcus, u4350841en_AU
local.description.embargo2099-12-31
local.description.notesImported from ARIESen_AU
local.description.refereedYes
local.identifier.absfor461105 - Reinforcement learningen_AU
local.identifier.ariespublicationa383154xPUB22397en_AU
local.identifier.doi10.48550/arXiv.2012.10200en_AU
local.identifier.thomsonID000681269800052
local.publisher.urlhttps://ojs.aaai.org/index.php/AAAI/article/view/17074/16881en_AU
local.type.statusPublished Versionen_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
17074-Article Text-20568-1-2-20210518.pdf
Size:
168.27 KB
Format:
Adobe Portable Document Format
Description: