Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

BenchMake: turn any scientific data set into a reproducible benchmark

dc.contributor.authorBarnard, A. S.en
dc.date.accessioned2025-12-23T08:40:32Z
dc.date.available2025-12-23T08:40:32Z
dc.date.issued2025-09-30en
dc.description.abstractBenchmark data sets are a cornerstone of machine learning development and applications, ensuring new methods are robust, reliable and competitive. The relative rarity of benchmark sets in computational science, due to the uniqueness of the problems and the pace of change in the associated domains, makes evaluating new innovations difficult for computational scientists. In this paper a new tool is developed and tested to potentially turn any of the increasing numbers of scientific data sets made openly available into a benchmark accessible to the community. BenchMake uses non-negative matrix factorization to deterministically identify and isolate challenging edge cases on the convex hull (the smallest convex set that contains all existing data instances) and partitions a required fraction of matched data instances into a testing set that maximizes divergence and statistical significance, across tabular, graph, image, signal and textual modalities. BenchMake splits are compared to establish splits and random splits using ten publicly available benchmark sets from different areas of science, with different sizes, shapes, distributions.en
dc.description.sponsorshipComputational resources for this project were supplied by the National Computing Infrastructure (NCI) [Grant Number p00]. ASB would like to thank Chloe Lin, Amanda Parker, Ben Mashford and Dan Andrews for beta testing.en
dc.description.statusPeer-revieweden
dc.format.extent11en
dc.identifier.otherWOS:001550096300001en
dc.identifier.otherORCID:/0000-0002-4784-2382/work/190976761en
dc.identifier.scopus105013207339en
dc.identifier.urihttps://hdl.handle.net/1885/733796934
dc.language.isoenen
dc.provenanceOriginal Content from this work may be used under the terms of the Creative Commons Attribution 4.0 licence. Any further distribution of this work must maintain attribution to the author(s) and the title of the work, journal citation and DOI.en
dc.rights© 2025 The Author(s). Published by IOP Publishing Ltd.en
dc.sourceMachine Learning: Science and Technologyen
dc.subjectArchetypal analysisen
dc.subjectBenchmarken
dc.subjectData spliten
dc.subjectModel evaluationen
dc.subjectSamplingen
dc.subjectScientific dataen
dc.titleBenchMake: turn any scientific data set into a reproducible benchmarken
dc.typeJournal articleen
dspace.entity.typePublicationen
local.bibliographicCitation.lastpage11en
local.bibliographicCitation.startpage1en
local.contributor.affiliationBarnard, A. S.; School of Computing, ANU College of Systems and Society, The Australian National Universityen
local.identifier.citationvolume6en
local.identifier.doi10.1088/2632-2153/adf810en
local.identifier.pure2f84279a-b0ea-46c1-864f-b14c97b3ae81en
local.identifier.urlhttps://www.scopus.com/pages/publications/105013207339en
local.type.statusPublisheden

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Barnard_2025_Mach._Learn._Sci._Technol._6_030502.pdf
Size:
458.72 KB
Format:
Adobe Portable Document Format