Daswani, MayankSunehag, PeterHutter, Marcus2015-08-142015-08-141532-4435http://hdl.handle.net/1885/14724There has recently been much interest in history-based methods using suffix trees to solve POMDPs. However, these suffix trees cannot efficiently represent environments that have long-term dependencies. We extend the recently introduced CTΦMDP algorithm to the space of looping suffix trees which have previously only been used in solving deterministic POMDPs. The resulting algorithm replicates results from CTΦMDP for environments with short term dependencies, while it outperforms LSTM-based methods on TMaze, a deep memory environment.© 2012 M. Daswani, P. Sunehag & M. Hutter. Author can archive publisher’s version/PDF. http://www.sherpa.ac.uk/romeo/issn/1532-4435/ as at 14/8/15looping suffix treesMarkov decision processreinforcement learningpartial observabilityMonte Carlo searchrational agentsFeature reinforcement learning using looping suffix trees2012-12