Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

NovPhy: A physical reasoning benchmark for open-world AI systems

dc.contributor.authorPinto, Vimukthinien
dc.contributor.authorGamage, Chathuraen
dc.contributor.authorXue, Chengen
dc.contributor.authorZhang, Pengen
dc.contributor.authorNikonova, Ekaterinaen
dc.contributor.authorStephenson, Matthewen
dc.contributor.authorRenz, Jochenen
dc.date.accessioned2025-05-23T08:22:08Z
dc.date.available2025-05-23T08:22:08Z
dc.date.issued2024en
dc.description.abstractDue to the emergence of AI systems that interact with the physical environment, there is an increased interest in incorporating physical reasoning capabilities into those AI systems. But is it enough to only have physical reasoning capabilities to operate in a real physical environment? In the real world, we constantly face novel situations we have not encountered before. As humans, we are competent at successfully adapting to those situations. Similarly, an agent needs to have the ability to function under the impact of novelties in order to properly operate in an open-world physical environment. To facilitate the development of such AI systems, we propose a new benchmark, NovPhy, that requires an agent to reason about physical scenarios in the presence of novelties and take actions accordingly. The benchmark consists of tasks that require agents to detect and adapt to novelties in physical scenarios. To create tasks in the benchmark, we develop eight novelties representing a diverse novelty space and apply them to five commonly encountered scenarios in a physical environment, related to applying forces and motions such as rolling, falling, and sliding of objects. According to our benchmark design, we evaluate two capabilities of an agent: the performance on a novelty when it is applied to different physical scenarios and the performance on a physical scenario when different novelties are applied to it. We conduct a thorough evaluation with human players, learning agents, and heuristic agents. Our evaluation shows that humans' performance is far beyond the agents' performance. Some agents, even with good normal task performance, perform significantly worse when there is a novelty, and the agents that can adapt to novelties typically adapt slower than humans. We promote the development of intelligent agents capable of performing at the human level or above when operating in open-world physical environments. Benchmark website: https://github.com/phy-q/novphy.en
dc.description.sponsorshipThis research was sponsored by the Defense Advanced Research Projects Agency (DARPA) and the Army Research Office (ARO) and was accomplished under Cooperative Agreement Number W911NF-20-2-0002. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the DARPA or ARO, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation herein.en
dc.description.statusPeer-revieweden
dc.identifier.issn0004-3702en
dc.identifier.otherORCID:/0000-0001-9667-623X/work/184099421en
dc.identifier.otherORCID:/0000-0003-3928-2255/work/184101974en
dc.identifier.scopus85201499939en
dc.identifier.urihttp://www.scopus.com/inward/record.url?scp=85201499939&partnerID=8YFLogxKen
dc.identifier.urihttps://hdl.handle.net/1885/733751813
dc.language.isoenen
dc.rightsPublisher Copyright: © 2024 The Authorsen
dc.sourceArtificial Intelligenceen
dc.subjectAI evaluationen
dc.subjectNovelty adaptationen
dc.subjectNovelty benchmarken
dc.subjectNovelty detectionen
dc.subjectOpen-world learningen
dc.subjectPhysical reasoningen
dc.titleNovPhy: A physical reasoning benchmark for open-world AI systemsen
dc.typeJournal articleen
dspace.entity.typePublicationen
local.contributor.affiliationPinto, Vimukthini; School of Computing, ANU College of Systems and Society, The Australian National Universityen
local.contributor.affiliationGamage, Chathura; School of Computing, ANU College of Systems and Society, The Australian National Universityen
local.contributor.affiliationXue, Cheng; School of Computing, ANU College of Systems and Society, The Australian National Universityen
local.contributor.affiliationZhang, Peng; School of Computing, ANU College of Systems and Society, The Australian National Universityen
local.contributor.affiliationNikonova, Ekaterina; Australian National Universityen
local.contributor.affiliationStephenson, Matthew; Flinders Universityen
local.contributor.affiliationRenz, Jochen; School of Computing, ANU College of Systems and Society, The Australian National Universityen
local.identifier.citationvolume336en
local.identifier.doi10.1016/j.artint.2024.104198en
local.identifier.pure8f8d00f6-c5ab-4c88-b05f-335d0781003ben
local.identifier.urlhttps://www.scopus.com/pages/publications/85201499939en
local.type.statusPublisheden

Downloads