Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Vision Language Model Guided Zero-shot Classification

Loading...
Thumbnail Image

Date

Authors

Yao, Haodong

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Among various core tasks in Computer Vision, 2D image and 3D object classification are fundamental tasks which serve as the foundation for numerous applications including scene understanding, robotics and autonomous navigation. Vision Language Model (VLMs) are deep learning architectures designed to process and understand both visual and textual information simultaneously. This thesis takes a close look at Vision-Language Models in classification tasks, with a particular emphasis on zero-shot settings in both 2D and 3D scenarios. We provide a comprehensive overview of Vision-Language Models, focusing on their pretraining datasets, architectural components, learning strategies, and representative models. By comparing with supervised 2D approaches including shell learning along with conventional 3D classification methods, in-depth experiments and analysis have been conducted from various perspectives, including classification performance, semantic clustering and computational efficiency.

Description

Keywords

Citation

Source

Book Title

Entity type

Access Statement

License Rights

Restricted until

Downloads

File
Description