Lattimore, TorHutter, Marcus2015-12-10October 8-9783319116617http://hdl.handle.net/1885/58180We consider a general reinforcement learning problem and show that carefully combining the Bayesian optimal policy and an exploring policy leads to minimax sample-complexity bounds in a very general class of (history-based) environments. We also prove lower bounds and show that the new algorithm displays adaptive behaviour when the environment is easier than worst-case.Copyright Information: © Springer International Publishing Switzerland 2014. http://www.sherpa.ac.uk/romeo/issn/0302-9743/..."Author's post-print on any open access repository after 12 months after publication" from SHERPA/RoMEO site (as at 13/08/15)Author/s retain copyrightBayesian reinforcement learning with exploration201410.1007/978-3-319-11662-4_132016-02-24