Niu, Xinlei2026-08-132026-08-13https://hdl.handle.net/1885/733814249Audio is an important modality through which humans perceive and interpret the world. From speech and environmental sounds to music, it contains linguistic, emotional, and contextual information that can represent artistic expression. With the rapid progress of machine learning and generative AI, generative audio synthesis has emerged as a powerful approach to automatically producing realistic and diverse sounds, speech, and music content. However, current methods have challenges in requiring large computational resources, lacking flexibility across audio content, and offering limited controllability to users in practice. This thesis aims to tackle these limitations by developing audio synthesis techniques that are efficient, flexible, and controllable. This thesis explores how generative models can be designed to operate effectively under training resource constraints, adapt across different and diverse creative audio content, and allow for fine-grained and customized control to users. To achieve these goals, six complementary methods are proposed in this thesis. Among them, SoundLoCD in Chapter 3 and BDP in Chapter 5 focus on training efficiency, introducing a lightweight diffusion framework and a discrete latent optimal path modeling framework that preserve quality while reducing training computational requirements and strategies. HybridVC in Chapter 6 and SoundMorpher in Chapter 7 advance flexibility by enabling flexible voice conversion and seamless sound morphing across multi-modalities conditions. BVS in Chapter 4 and SteerMusic in Chapter 8 enhance synthesizing controllability, which allows fine-grained alignment between audio and visual context and empowers user-driven editing in music generation. In summary, these contributions form a unified perspective on AI-generated content of audio synthesis, which treats speech, sound, and music not as isolated problems but as interconnected domains of users immersive auditory experience. The resulting proposed methods move toward a future where audio generation is accessible, adaptive, and deeply integrated into creative, assistive, and interactive applications.en-AUGenerative Audio Synthesis in Sound, Speech, and Music2026