Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Learning-integrated Methods for Spatial Audio Capture and Reproduction

Loading...
Thumbnail Image

Date

Authors

Chen, Xingyu

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

This thesis presents a study on advancing spatial audio technologies using deep learning methods, with a focus on three major areas: sound field estimation, Head-Related Transfer Function(HRTF) interpolation, and monaural speech enhancement on drones. Spatial audio is integral to immersive applications such as entertainment, teleconferencing, and AR/VR applications including gaming and defense applications. However, challenges persist due to technological constraints like the limited spatial resolution of microphone arrays, which impede the full capture and reproduction of sound scenes. As a result, there is a notable drive to enhance spatial audio technologies to improve user experience. The integration of deep learning in spatial audio is transformative, paralleling its impact in fields like image processing and computer vision. Deep learning's capability to manage complex, non-linear problems has started to revolutionize spatial audio, though its application in this field is still in the early stages of development. This thesis introduces innovative learning-based methods to address challenges in spatial audio. It employs Physics-Informed Neural Networks for sound field estimation to achieve accurate sound pressure estimation. For HRTF interpolation, it presents a spherical convolutional neural network designed to enhance interpolation accuracy. Furthermore, the thesis explores the application of transfer learning to speech enhancement on drones, focusing on improving clarity and reducing drone noise. The findings of this thesis not only contribute to the theoretical advancements in the field of spatial audio but also have practical implications for enhancing user experience in many immersive audio applications including those using VR/AR techniques. It also enables future use of drone embedded microphone arrays for spatial audio capture This work contributes insights into the application of deep learning technologies in spatial audio and offers ideas for future studies to build on and improve these methods.

Description

Keywords

Citation

Source

Book Title

Entity type

Access Statement

License Rights

Restricted until

2025-05-14

Downloads

File
Description