Hutter, Marcus2015-09-022015-09-02978-3-540-43836-60302-9743http://hdl.handle.net/1885/15096© Springer-Verlag Berlin Heidelberg 2002. http://www.sherpa.ac.uk/romeo/issn/0302-9743/..."Author's post-print on any open access repository after 12 months after publication" from SHERPA/RoMEO site (as at 2/09/15)Rational agentssequential decision theoryreinforcement learningvalue functionBayes mixturesself-optimizing policiesPareto-optimalityunbounded effective horizon(non) Markov decision processesSelf-optimizing and Pareto-optimal policies in general environments based on Bayes-mixtures200210.1007/3-540-45435-7_25