Exact Reduction of Huge Action Spaces in General Reinforcement Learning
| dc.contributor.author | Majeed, Sultan | |
| dc.contributor.author | Hutter, Marcus | |
| dc.coverage.spatial | virtual | |
| dc.date.accessioned | 2024-01-29T23:01:34Z | |
| dc.date.created | February 2–9, 2021 | |
| dc.date.issued | 2021 | |
| dc.date.updated | 2022-10-02T07:18:42Z | |
| dc.description.abstract | The reinforcement learning (RL) framework formalizes the notion of learning with interactions. Many real-world problems have large state-spaces and/or action-spaces such as in Go, StarCraft, protein folding, and robotics or are non-Markovian, which cause significant challenges to RL algorithms. In this work we address the large action-space problem by sequentializing actions, which can reduce the action-space size significantly, even down to two actions at the expense of an increased planning horizon. We provide explicit and exact constructions and equivalence proofs for all quantities of interest for arbitrary history-based processes. In the case of MDPs, this could help RL algorithms that bootstrap. In this work we show how action-binarization in the nonMDP case can significantly improve Extreme State Aggregation (ESA) bounds. ESA allows casting any (non-MDP, non-ergodic, history-based) RL problem into a fixed-sized non-Markovian state-space with the help of a surrogate Markovian process. On the upside, ESA enjoys similar optimality guarantees as Markovian models do. But a downside is that the size of the aggregated state-space becomes exponential in the size of the action-space. In this work, we patch this issue by binarizing the action-space. We provide an upper bound on the number of states of this binarized ESA that is logarithmic in the original action-space size, a double-exponential improvement. | en_AU |
| dc.description.sponsorship | This work has been supported by Australian Research Council grant DP150104590 | en_AU |
| dc.format.mimetype | application/pdf | en_AU |
| dc.identifier.isbn | 978-1-57735-866-4 | en_AU |
| dc.identifier.uri | http://hdl.handle.net/1885/312404 | |
| dc.language.iso | en_AU | en_AU |
| dc.publisher | The AAAI Press | en_AU |
| dc.relation | http://purl.org/au-research/grants/arc/DP150104590 | en_AU |
| dc.relation.ispartofseries | THIRTY-FIFTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE | en_AU |
| dc.rights | © 2021 Association for the Advancement of Artificial Intelligence | en_AU |
| dc.source.uri | https://ojs.aaai.org/index.php/AAAI/article/view/17074/16881 | en_AU |
| dc.title | Exact Reduction of Huge Action Spaces in General Reinforcement Learning | en_AU |
| dc.type | Conference paper | en_AU |
| dcterms.accessRights | Free Access via publisher website | en_AU |
| local.bibliographicCitation.lastpage | 8883 | en_AU |
| local.bibliographicCitation.startpage | 8874 | en_AU |
| local.contributor.affiliation | Majeed, Sultan, College of Engineering and Computer Science, ANU | en_AU |
| local.contributor.affiliation | Hutter, Marcus, College of Engineering and Computer Science, ANU | en_AU |
| local.contributor.authoruid | Majeed, Sultan, u5447242 | en_AU |
| local.contributor.authoruid | Hutter, Marcus, u4350841 | en_AU |
| local.description.embargo | 2099-12-31 | |
| local.description.notes | Imported from ARIES | en_AU |
| local.description.refereed | Yes | |
| local.identifier.absfor | 461105 - Reinforcement learning | en_AU |
| local.identifier.ariespublication | a383154xPUB22397 | en_AU |
| local.identifier.doi | 10.48550/arXiv.2012.10200 | en_AU |
| local.identifier.thomsonID | 000681269800052 | |
| local.publisher.url | https://ojs.aaai.org/index.php/AAAI/article/view/17074/16881 | en_AU |
| local.type.status | Published Version | en_AU |
Downloads
Original bundle
1 - 1 of 1
Loading...
- Name:
- 17074-Article Text-20568-1-2-20210518.pdf
- Size:
- 168.27 KB
- Format:
- Adobe Portable Document Format
- Description: