PAC bounds for discounted MDPs

Lattimore, Tor; Hutter, Marcus

doi:10.1007/978-3-642-34106-9_26

A change is coming. Click to see a sneak peek of the new Open Research Repository.

PAC bounds for discounted MDPs

link to publisher version

Altmetric Citations

Lattimore, Tor; Hutter, Marcus

Description

We study upper and lower bounds on the sample-complexity of learning near-optimal behaviour in finite-state discounted Markov Decision Processes (mdps). We prove a new bound for a modified version of Upper Confidence Reinforcement Learning (ucrl) with only cubic dependence on the horizon. The bound is unimprovable in all parameters except the size of the state/action space, where it depends linearly on the number of non-zero transition probabilities. The lower bound strengthens previous work by...[Show more] being both more general (it applies to all policies) and tighter. The upper and lower bounds match up to logarithmic factors provided the transition matrix is not too dense.

Collections	ANU Research Publications
Date published:	2012
Type:	Conference paper
URI:	http://hdl.handle.net/1885/69046
Source:	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
DOI:	10.1007/978-3-642-34106-9_26

Download

There are no files associated with this item.

Show full item record

PAC bounds for discounted MDPs

Altmetric Citations

Description

Download