Learning Deep Features for Limited-data Visual Problems
Abstract
In the realm of deep computer vision, especially when faced with the challenge of limited data, leveraging innovative mechanisms for learning deep features becomes critical. This thesis presents five works to tackle limited-data visual problems, including low-shot and open-set scenarios.
In the first work, few-shot recognition is addressed by introducing an attention agent to the backbone network. This agent, trained via reinforcement learning, adaptively locates representative regions on feature maps of the backbone. The policy gradient algorithm aids this process, focusing the agent mechanism for better generalization across unknown classes. The results demonstrate progressively discriminative feature maps, evident in the performance on few-shot image classification.
The problem of audio-visual zero-shot learning is studied through a hyperbolic learning approach in the second work. With a recognized ability to discover hierarchical structure in data, the hyperbolic projection facilitates learning discriminative embedding features for audio-visual samples. Moreover, cross-modality alignment in the hyperbolic space is key in this approach, leading to clear improvements in several benchmark datasets.
Moving to the challenge of open-set recognition, curved space embeddings continue to take the spotlight in the third work. The previous work demonstrates that hyperbolic geometry can encode richer structural information. For the open-set recognition task, we extend the approach using more types of curved geometries, including spherical and mixed geometries. These geometric constraints are interchangeably employed depending on the task. Two innovative geometric modules are proposed, replacing original Euclidean classifiers and improving performance in many visual open-set recognition tasks.
In the third work, we find that open-set semantic segmentation (OSS) sometimes leaves big regions on the image unprocessed. To this end, in the fourth work, we extend OSS to the more encompassing task, \ie, generalized open-set semantic segmentation (GOSS). GOSS not only detects unknown regions but also enhances the agent's decision-making ability by further predicting more segments inside the unknown region. By integrating both pixel classification and clustering mechanisms, GOSS offers a more comprehensive open-set segmentation solution.
Finally, in the fifth work, the thesis sheds light on the practical challenges of point cloud learning. A novel Point Cut-and-mix (PointCaM) mechanism is introduced to address open-set point cloud learning scenarios. PointCaM mechanism simulates out-of-distribution data and distinguishes between known and unknown data, showcasing promising results in its experimental evaluation.
In conclusion, throughout this thesis, five innovative techniques or strategies are presented to optimize the learning of deep features, particularly in scenarios marked by data scarcity.
Description
Keywords
Citation
Collections
Source
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
2024-06-26
Downloads
File
Description
Thesis Material