Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Characterising Algorithm Debt in Machine and Deep Learning Systems

Loading...
Thumbnail Image

Date

Authors

Simon, Emmanuel Iko Ojo

Journal Title

Journal ISSN

Volume Title

Publisher

IEEE Computer Society

Access Statement

Research Projects

Organizational Units

Journal Issue

Abstract

Technical Debt (TD) refers to the long-term costs of suboptimal choices made for short-term gains. Algorithm Debt (AD), a type of TD, refers to the sub-optimal implementation of an algorithm, which can negatively impact software systems as they evolve. Machine and Deep Learning (ML/DL) systems are prone to AD due to the complexity of the methods and data dependencies. AD in ML/DL systems leads to model degradation and poor scalability over time. However, there are no automated tools for its identification and the causes, effects, and mitigation strategies remain largely unknown, creating a knowledge gap. To address this gap, the aim of this research is to investigate the causes, effects, and mitigation strategies of AD in ML/DL systems. Using a mixed-methods approach, we conducted a review of 35 primary studies. We also performed experiments with ML/DL models for AD identification, and conducted interviews with 21 practitioners, complimented by an ongoing online questionnaire. Preliminary Findings suggest that the causes of AD can be grouped into issues related to data quality, ML knowledge gaps, and challenges in model training. From the experiments, the approach of the training a Logistic Regression model with AD custom features achieved the highest F1 score of 54%, outper-forming other techniques. This study contributes to software engineering by proposing a framework that highlights the causes, effects, and mitigation strategies for AD in ML/DL systems. The framework serves as a guideline for practitioners and offers insights for developing automated tools and effective management strategies for AD in ML/DL systems.

Description

Citation

Source

Book Title

Proceedings - 2025 IEEE/ACM 47th International Conference on Software Engineering, ICSE-Companion 2025

Entity type

Publication

Access Statement

License Rights

Restricted until