Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Localizing Objects with Weak Supervision

dc.contributor.authorSun, Weixuan
dc.date.accessioned2024-01-17T14:43:44Z
dc.date.available2024-01-17T14:43:44Z
dc.date.issued2024
dc.description.abstractIn the topic of object localization, the process involves identifying objects, determining their locations, and estimating their extents. This topic plays a vital role in computer vision and has multiple applications. Many existing localization systems rely on techniques like semantic segmentation or object detection models, which typically require exhaustively labelled data. However, acquiring fully-labelled data can be expensive and time-consuming. To overcome this challenge, several alternative weak labels such as class labels, audio waves, and text captions, can be utilized. These weak labels are either simple to gather from sources like the internet or specialized equipment, or they are less costly to annotate. In this thesis, our aim is to derive object localization using weak labels. The contributions are three-folded. The first contribution includes three successive works that localize precise object masks from class labels for image-level WSSS. 1) Existing image-level WSSS approaches necessitate a complex training pipeline to refine object localization. However, we investigate the class activation mechanism of CNN networks and find that object activation can be greatly improved without retraining the baseline model. We present an inference-only solution named class conditional inference module which reveals the classification network's hidden object activation to generate more integral object localization. 2) Most existing methods rely on CNN networks to generate class-wise activation, while we find that simply transplanting a standard CNN-based localization method to transformers results in extensive noise. To overcome this issue and leverage the advantages of transformer networks, we offer a new transformer-based class-activation method which uses gradients flowing back through the self-attention modules. We demonstrate that it can produce reliable object localization, and incorporate it into an end-to-end WSSS framework. 3) To further incorporate the capability of transformer networks that encode pair-wise links between image areas, we propose all-pairs consistency regularization (ACR), which is a self-supervised training regularization, and enables the transformers to capture more accurate object localization. The aforementioned three works are compared with existing image-level WSSS methods, evaluating on standard segmentation datasets. As the second contribution, we introduce a novel approach for comprehending indoor scenes by utilizing sparse bounding boxes and 3D data. First, objects are localized as point clouds in 3D space and then are projected back to 2D images to obtain object localization masks. Further, our method works in a recursive manner to gradually refine the localization. This method is assessed against previous bounding box-based WSSS algorithms using a public indoor scene dataset. In our third contribution, we investigate audio-visual source localization (AVSL), where the goal is to locate objects based on audio waves. We notice that previous approaches employ a contrastive learning strategy but disregard the issue of false negatives. To address this issue, we propose false negative aware contrastive (FNAC) learning framework. We suggest two new regularization terms during the contrastive learning, where we detect potential false negatives and mitigate their impact. The efficacy of this method is demonstrated on several AVSL benchmarks. In conclusion, we investigate the problem of locating objects in images using weak labels. We illustrate the efficacy of our methods by comparing them to the current literatures and evaluating them on the public benchmarks. Our contributions bridge the gap between weak labels and pixel-level predictions, attempting to overcome the limitations of requiring fully annotated data for object localization. By this thesis, we strive to achieve a better understanding of building effective and efficient systems for localization of objects using weak labels.
dc.identifier.urihttp://hdl.handle.net/1885/311579
dc.language.isoen_AU
dc.titleLocalizing Objects with Weak Supervision
dc.typeThesis (PhD)
local.contributor.supervisorBarnes, Nicholas
local.identifier.doi10.25911/9VN3-0T42
local.mintdoimint
local.thesisANUonly.author928d7d6e-156f-4f5d-ad2c-e6617369328c
local.thesisANUonly.keycb2f4047-ffd0-9225-dc08-b4631396e30f
local.thesisANUonly.title000000024561_TC_1

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
weixuan_thesis_20240214.pdf
Size:
34.16 MB
Format:
Adobe Portable Document Format
Description:
Thesis Material