Partially Observable Reference Policy Programming
Loading...
Date
Authors
Kim, Edward
Kurniawati, Hanna
Journal Title
Journal ISSN
Volume Title
Publisher
ijcai.org
Access Statement
Abstract
This paper proposes Partially Observable Reference Policy Programming, a novel anytime online
approximate POMDP solver which samples meaningful future histories very deeply while simultaneously forcing a gradual policy update. We provide theoretical guarantees for the algorithm’s underlying scheme which say that the performance
loss is bounded by the average of the sampling approximation errors rather than the usual maximum;
a crucial requirement given the sampling sparsity
of online planning. Empirical evaluations on two
large-scale problems with dynamically evolving
environments—including a helicopter emergency
scenario in the Corsica region requiring approximately 150 planning steps—corroborate the theoretical results and indicate that our solver considerably outperforms current online benchmarks.
Description
Keywords
Citation
Collections
Source
Type
Book Title
Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJCAI 2025, Montreal, Canada, August 16-22, 2025
Entity type
Publication