Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Optimal load balancing techniques for block-cyclic decompositions for matrix factorization

dc.contributor.authorStrazdins, Peteren_US
dc.date.accessioned2003-07-03en_US
dc.date.accessioned2004-05-19T12:25:55Zen_US
dc.date.accessioned2011-01-05T08:37:59Z
dc.date.available2004-05-19T12:25:55Zen_US
dc.date.available2011-01-05T08:37:59Z
dc.date.created1998en_US
dc.date.issued1998en_US
dc.description.abstractIn this paper, we present a new load balancing technique, called panel scattering, which is generally applicable for parallel block-partitioned dense linear algebra algorithms, such as matrix factorization. Here, the panels formed in such computation are divided across their length, and evenly (re-)distributed among all processors. It is shown how this technique can be eÆciently implemented for the general block-cyclic matrix distribution, requiring only the collective communication primitives that required for block-cyclic parallel BLAS. In most situations, panel scattering yields optimal load balance and cell computation speed across all stages of the computation. It has also advantages in naturally yielding good memory access patterns. Compared with traditional methods which minimize communication costs at the expense of load balance, it has a small (in some situations negative) increase in communication volume costs. It however incurs extra communication startup costs, but only by a factor not exceeding 2. To maximize load balance and minimize the cost of panel re-distribution, storage block sizes should be kept small; furthermore, in many situations of interest, there will be no significant communication startup penalty for doing so. Results will be given on the Fujitsu AP+ parallel computer, which will compare the performance of panel scattering with previously established methods, for LU, LLT and QR factorization. These are consistent with a detailed performance model for LU factorization for each method that is developed here.en_US
dc.format.extent338101 bytesen_US
dc.format.extent356 bytesen_US
dc.format.mimetypeapplication/pdfen_US
dc.format.mimetypeapplication/octet-streamen_US
dc.identifier.urihttp://hdl.handle.net/1885/40738en_US
dc.identifier.urihttp://digitalcollections.anu.edu.au/handle/1885/40738
dc.language.isoen_AUen_US
dc.subjectdense linear algebraen_US
dc.subjectblock cyclic decompositionen_US
dc.subjectstorage blockingen_US
dc.subjectalgorithmic blockingen_US
dc.subjectphysically based matrix distributionen_US
dc.subjectTR-CSen_US
dc.titleOptimal load balancing techniques for block-cyclic decompositions for matrix factorizationen_US
dc.typeWorking/Technical Paperen_US
local.citationTR-CS-98-10en_US
local.contributor.affiliationDepartment of Computer Science, FEITen_US
local.contributor.affiliationANUen_US
local.description.refereednoen_US
local.identifier.citationmonthsepen_US
local.identifier.citationyear1998en_US
local.identifier.eprintid1559en_US
local.rights.ispublishedyesen_US

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
TR-CS-98-10.pdf
Size:
330.18 KB
Format:
Adobe Portable Document Format