Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Visual Recognition From Structured Supervision

Loading...
Thumbnail Image

Date

Authors

Fonseca De Santa Cruz Oliveira, Rodrigo

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Visual recognition of semantically meaningful entities like objects, actions, and poses in images and videos is a long standing goal of computer vision. In the last decades, we have seen progress towards this goal with the development of machine learning models that leverage huge volumes of human annotated data to perform very accurate recognition of a predefined set of visual entities. However, moving forward, this approach presents significant limitations since annotated datasets are expensive to collect, only contemplate a small fraction of the real world, and the labelling task itself is prone to inconsistency and ambiguity on denoting visual entities. Therefore, this reliance on exhaustive labeling is indeed the key obstacle to the fulfillment of such a goal. In this thesis, we propose methods that reduce the need for human supervision by leveraging the structure in the visual world targeting visual recognition in difficult scenarios where annotated data is scarce and the visual concepts are innumerable or ambiguous. We call this approach structured supervised learning and explore three instances of structured supervision. We start by exploring structure in the output of visual recognition models to learn better models for ranking images according to a predefined criteria, like the visual attribute "smiling". Towards this end, we first cast the problem of image ranking as the problem of predicting the correct permutation of a set of images. Then, we leverage the geometrical structure of permutation matrices in order to learn accurate image rankers. Next, we explore the self-supervision that can be extracted from input visual data itself. More specifically, unlabeled visual data itself encompasses rich spatial (and temporal) structure that can be explored in order to learn representations useful for generic visual recognition tasks. In contrast to human annotators, this form of self-supervision is cheap and abundant. Following this idea, we use the spatial layout of objects as a supervisory signal to learn transferable image representations from unlabeled data for object recognition tasks such as image classification, object detection, and object segmentation. Last, we observe that the visual world is fundamentally compositional and complex visual concepts are structured compositions of simple primitive concepts. We build in this insight and formulate frameworks to unambiguously describe and recognize compositional visual concepts in images and videos by exploring structural information in model space. More specifically, we classify objects from boolean expressions of object attributes and infer activities from regular expressions of atomic actions. The proposed models can predict unseen, subcategories and specific instances of complex visual concepts without any additional annotation effort, resulting in a more feasible direction to fulfill the visual recognition goal.

Description

Keywords

Citation

Source

Book Title

Entity type

Access Statement

License Rights

Restricted until

Downloads