Toward Robust Visual Perception for Autonomous Systems: Advancements in Cross-View Localisation and Temporal Fusion
| dc.contributor.author | Wang, Shan | |
| dc.date.accessioned | 2026-04-29T12:58:55Z | |
| dc.date.available | 2026-04-29T12:58:55Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Autonomous driving systems require robust visual perception to operate safely in diverse and dynamic environments. This thesis advances visual perception by focusing on two key areas: cross-view localisation and temporal fusion, aiming to reduce reliance on expensive sensors and pre-built 3D maps while achieving accurate and reliable performance. To overcome the limitations of traditional localisation methods, which often depend on costly survey-grade mapping vehicles, we propose a scalable alternative that leverages high-resolution satellite imagery as a global map. By aligning ground-view images from vehicle-mounted cameras with aerial views, we achieve accurate cross-view localisation using only visual inputs. We introduce three novel localisation solutions: (1) A visual-LiDAR hybrid method that uses 3D LiDAR points to establish correspondences between views, featuring a geometric-aligned feature extractor and iterative pose refinement. (2) A purely visual method that detects view-consistent ground keypoints and exploits spatial camera constraints to suppress dynamic and seasonal noise. (3) An enhanced visual method that aggregates off-ground aerial features onto ground-level pixels via orthogonal-view transformation, improving alignment and orientation estimation. Beyond localisation, we tackle the challenge of occluded lane markings through a homography-guided temporal fusion module. By combining information from adjacent video frames with a surface-normal estimator, this approach improves lane segmentation under visual occlusions while remaining computationally efficient. Together, these contributions form a robust visual perception framework. By reducing dependency on external sensors and pre-built maps, our methods enable autonomous systems to achieve accurate, reliable, and scalable perception in real-world conditions. | |
| dc.identifier.uri | https://hdl.handle.net/1885/733808736 | |
| dc.language.iso | en_AU | |
| dc.title | Toward Robust Visual Perception for Autonomous Systems: Advancements in Cross-View Localisation and Temporal Fusion | |
| dc.type | Thesis (PhD) | |
| local.contributor.affiliation | College of Systems and Society, The Australian National University | |
| local.contributor.supervisor | Nguyen, Vinh Chuong | |
| local.identifier.doi | 10.25911/Z9H4-XQ61 | |
| local.identifier.proquest | Yes | |
| local.identifier.researcherID | JVE-2291-2024 | |
| local.mintdoi | mint | |
| local.thesisANUonly.author | 5a9f329e-3ed9-49a4-841d-0b50ac252ce2 | |
| local.thesisANUonly.key | c81adad0-8baa-0f2a-836e-22c2c1c35278 | |
| local.thesisANUonly.title | 000000026822_TC_1 |
Downloads
Original bundle
1 - 1 of 1
Loading...
- Name:
- Wang_Thesis_2026.pdf
- Size:
- 61.63 MB
- Format:
- Adobe Portable Document Format
- Description:
- Thesis Material