Near-optimal PAC bounds for discounted MDPs
dc.contributor.author | Lattimore, Tor | |
dc.contributor.author | Hutter, Marcus | |
dc.date.accessioned | 2015-12-10T22:43:56Z | |
dc.date.issued | 2014 | |
dc.date.updated | 2016-02-24T10:33:52Z | |
dc.description.abstract | We study upper and lower bounds on the sample-complexity of learning near-optimal behaviour in finite-state discounted Markov Decision Processes (mdps). We prove a new bound for a modified version of Upper Confidence Reinforcement Learning (ucrl) with only cubic dependence on the horizon. The bound is unimprovable in all parameters except the size of the state/action space, where it depends linearly on the number of non-zero transition probabilities. The lower bound strengthens previous work by being both more general (it applies to all policies) and tighter. The upper and lower bounds match up to logarithmic factors provided the transition matrix is not too dense. | |
dc.identifier.issn | 0304-3975 | |
dc.identifier.uri | http://hdl.handle.net/1885/58388 | |
dc.publisher | Elsevier | |
dc.rights | Copyright Information: © 2014 Elsevier B.V. http://www.sherpa.ac.uk/romeo/issn/0304-3975/..."Author's post-print on open access repository after an embargo period of between 12 months and 48 months" from SHERPA/RoMEO site (as at 10/08/15) | |
dc.source | Theoretical Computer Science | |
dc.title | Near-optimal PAC bounds for discounted MDPs | |
dc.type | Journal article | |
local.bibliographicCitation.lastpage | 143 | |
local.bibliographicCitation.startpage | 125 | |
local.contributor.affiliation | Lattimore, Tor, University of Alberta | |
local.contributor.affiliation | Hutter, Marcus, College of Engineering and Computer Science, ANU | |
local.contributor.authoremail | u4350841@anu.edu.au | |
local.contributor.authoruid | Hutter, Marcus, u4350841 | |
local.description.embargo | 2037-12-31 | |
local.description.notes | Imported from ARIES | |
local.identifier.absfor | 080100 - ARTIFICIAL INTELLIGENCE AND IMAGE PROCESSING | |
local.identifier.absseo | 970108 - Expanding Knowledge in the Information and Computing Sciences | |
local.identifier.ariespublication | u4056230xPUB440 | |
local.identifier.citationvolume | 558 | |
local.identifier.doi | 10.1016/j.tcs.2014.09.029 | |
local.identifier.scopusID | 2-s2.0-84926305366 | |
local.identifier.uidSubmittedBy | u4056230 | |
local.type.status | Published Version |
Downloads
Original bundle
1 - 1 of 1
No Thumbnail Available
- Name:
- 01_Lattimore_Near-optimal_PAC_bounds_for_2014.pdf
- Size:
- 468.63 KB
- Format:
- Adobe Portable Document Format