Designing Resilient Parallel Multigrid Methods on Unstructured Adaptive Meshes
Abstract
In this thesis, we develop a class of fault-tolerant geometric multigrid method to address the loss of numerical data on faulty processors. Prior studies of fault-tolerant multigrid methods have focused on structured grids. In this research, the challenge lies in the fact that unstructured grids distributed across multiple processors may manifest as local hierarchical grids with unaligned boundaries. This unalignment can cause divergence in the local multigrid recovery process. We address the recovery convergence issue from two distinct angles. Firstly, we adjust the local hierarchical grids to align their boundaries. This approach reduces the problem to the standard multigrid situation, hence yields the convergence. The second approach takes the local unaligned grids as they are, and propose a novel algorithm that modifies the multigrid V-cycle algorithm based on energy control criterion. We prove the convergence of the new algorithm by using smoothing and approximation properties defined in nonnested subspaces. Additionally, we introduce a near-boundary treatment and a local energy criterion to improve the convergence rate and shorten the runtime of local recovery, respectively. Our numerical experiments confirm that the proposed algorithms restore the convergence of the multigrid V-cycle on unstructured grids with unaligned boundaries, while the algorithm agrees with the standard multigrid V-cycle on local grids with aligned boundaries. They are also shown to mitigate the fault impact and reduce the delays in the post-fault global multigrid iterations completely. To automate the recovery process and to optimise its efficiency, we also introduce adaptive stopping criteria derived from the discrete maximum principle that terminate the recovery process based on the characteristics of the grid and the underlying solution. Furthermore, we implement an asynchronous recovery strategy that utilises the healthy processors to continue computation while the faulty processor recovers. This strategy improves the overall recovery efficiency by reducing the time-to-solution of the recovery process.
Description
Keywords
Citation
Collections
Source
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
Downloads
File
Description
Thesis Material