Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication

dc.contributor.authorXia, Yufanen
dc.contributor.authorPierre, Marco De Laen
dc.contributor.authorBarnard, Amanda S.en
dc.contributor.authorBarca, Giuseppe Maria Junioren
dc.date.accessioned2026-03-20T14:41:36Z
dc.date.available2026-03-20T14:41:36Z
dc.date.issued2026en
dc.description.abstractThe GEneral Matrix Multiplication (GEMM) is one of the essential algorithms in scientific computing. Single-thread GEMM implementations are well-optimised with techniques like blocking and autotuning. However, due to the complexity of modern multi-core shared memory systems, it is challenging to determine the number of threads that minimises the multi-thread GEMM runtime. We present a proof-of-concept approach to building an Architecture and Data-Structure Aware Linear Algebra (ADSALA) software library that uses machine learning to optimise the runtime performance of BLAS routines. More specifically, our method uses a machine learning model on-the-fly to automatically select the optimal number of threads for a given GEMM task based on the collected training data. Test results on two different HPC node architectures, one based on a two-socket Intel Cascade Lake and the other on a two-socket AMD Zen 3, revealed a 25 to 40 per cent speedup compared to traditional GEMM implementations in BLAS when using GEMM of memory usage within 100 MB.en
dc.description.statusNot peer-revieweden
dc.format.extent11en
dc.identifier.otherdblp:journals/corr/abs-2601-09114en
dc.identifier.urihttps://hdl.handle.net/1885/733807519
dc.language.isoenen
dc.rightsDBLP License: DBLP's bibliographic metadata records provided through http://dblp.org/ are distributed under a Creative Commons CC0 1.0 Universal Public Domain Dedication. Although the bibliographic metadata records are provided consistent with CC0 1.0 Dedication, the content described by the metadata records is not. Content may be subject to copyright, rights of privacy, rights of publicity and other restrictions.en
dc.sourceCoRRen
dc.titleA Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplicationen
dc.typeJournal articleen
dspace.entity.typePublicationen
local.contributor.affiliationBarnard, Amanda S.; School of Computing, ANU College of Systems and Society, The Australian National Universityen
local.contributor.affiliationBarca, Giuseppe Maria Junior; School of Computing, ANU College of Systems and Society, The Australian National Universityen
local.identifier.doi10.48550/arXiv.2601.09114en
local.identifier.pure881e5241-418e-4dbb-879e-2c2c5157b88aen
local.type.statusPublisheden

Downloads