Hutter, Marcus2015-12-07October 7-3540466495http://hdl.handle.net/1885/28236Consider an agent interacting with an environment in cycles. In every interaction cycle the agent is rewarded for its performance. We compare the average reward U from cycle 1 to m (average value) with the future discounted reward V from cycle k to ∞ (dCopyright Information: © Springer-Verlag Berlin Heidelberg 2006. http://www.sherpa.ac.uk/romeo/issn/0302-9743/..."Author's post-print on any open access repository after 12 months after publication" from SHERPA/RoMEO site (as at 31/08/15)Keywords: Artificial intelligence; Asymptotic stability; Computer simulation; Interactive computer systems; Arbitrary reward sequences; Interaction cycle; Intelligent agentsGeneral discounting versus average reward20062016-02-24