Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Reward Potentials for Planning with Learned Neural Network Transition Models

dc.contributor.authorSay, Buser
dc.contributor.authorSanner, Scott
dc.contributor.authorThiebaux, Sylvie
dc.contributor.editorSchiex, T
dc.contributor.editorde Givry, S
dc.coverage.spatialStamford, United States
dc.date.accessioned2024-01-17T23:56:32Z
dc.date.createdSep 30 - Oct 4 2019
dc.date.issued2019
dc.date.updated2022-10-02T07:16:49Z
dc.description.abstractOptimal planning with respect to learned neural network (NN) models in continuous action and state spaces using mixed-integer linear programming (MILP) is a challenging task for branch-and-bound solvers due to the poor linear relaxation of the underlying MILP model. For a given set of features, potential heuristics provide an efficient framework for computing bounds on cost (reward) functions. In this paper, we model the problem of finding optimal potential bounds for learned NN models as a bilevel program, and solve it using a novel finite-time constraint generation algorithm. We then strengthen the linear relaxation of the underlying MILP model by introducing constraints to bound the reward function based on the precomputed reward potentials. Experimentally, we show that our algorithm efficiently computes reward potentials for learned NN models, and that the overhead of computing reward potentials is justified by the overall strengthening of the underlying MILP model for the task of planning over long horizons.en_AU
dc.format.mimetypeapplication/pdfen_AU
dc.identifier.isbn978-3-030-30047-0en_AU
dc.identifier.urihttp://hdl.handle.net/1885/311593
dc.language.isoen_AUen_AU
dc.publisherSpringeren_AU
dc.relation.ispartofseries25rd International Conference on the Principles and Practice of Constraint Programming, CP 2019en_AU
dc.rights© Springer Nature Switzerland AG 2019en_AU
dc.subjectNeural networksen_AU
dc.subjectPotential heuristicsen_AU
dc.subjectPlanningen_AU
dc.subjectConstraint generationen_AU
dc.titleReward Potentials for Planning with Learned Neural Network Transition Modelsen_AU
dc.typeConference paperen_AU
local.bibliographicCitation.lastpage689en_AU
local.bibliographicCitation.startpage674en_AU
local.contributor.affiliationSay, Buser, University of Torontoen_AU
local.contributor.affiliationSanner, Scott, University of Toronto, Canadaen_AU
local.contributor.affiliationThiebaux, Sylvie, College of Engineering and Computer Science, ANUen_AU
local.contributor.authoruidThiebaux, Sylvie, u4033066en_AU
local.description.embargo2099-12-31
local.description.notesImported from ARIESen_AU
local.description.refereedYes
local.identifier.absfor460209 - Planning and decision makingen_AU
local.identifier.ariespublicationa383154xPUB11880en_AU
local.identifier.doi10.1007/978-3-030-30048-7_39en_AU
local.identifier.scopusID2-s2.0-85075746345
local.identifier.thomsonIDWOS:000560404200039
local.publisher.urlhttps://link.springer.com/en_AU
local.type.statusPublished Versionen_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
978-3-030-30048-7_39.pdf
Size:
1.23 MB
Format:
Adobe Portable Document Format
Description: