Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Molecular and polymer representations for machine learning

Loading...
Thumbnail Image

Authors

Lin, Chloe
Taylor, John A.
Barnard, Amanda S.
Connal, Luke A.
Pollard, Brett Leslie
Parker, Amanda J.

Journal Title

Journal ISSN

Volume Title

Publisher

Access Statement

Research Projects

Organizational Units

Journal Issue

Abstract

The success of AI-driven materials discovery ultimately depends on how we represent molecular structure. Traditional approaches based on expert-defined descriptors, fingerprints, and symbolic notations have enabled advances in property prediction and molecular design. These often struggle, however, to generalise across structurally complex chemical systems. This issue is particularly striking when structural hierarchy and stochasticity are not explicitly encoded or when datasets lack access to these higher-order features. Graph neural networks and large language models offer new pathways in representation learning, enabling molecular and polymer representations to be learned directly from structural inputs. These machine-learned embeddings promise more expressive, scalable, and transferable representations, potentially transforming polymer informatics and materials discovery. In this review, we survey the evolution of molecular representations from traditional descriptors to modern embedding-based approaches, examine how these approaches have been applied to small molecules and adapted for polymer systems, and analyse the challenges posed by stochasticity, structural hierarchy and limited data availability in polymer systems. We argue that future progress will depend not only on improved model architectures but also on the development of better-curated datasets and shared benchmarks, particularly for polymers, where stochasticity and structural hierarchy complicate both representation and evaluation. Advancing representation learning in chemistry will therefore require the deliberate integration of domain knowledge, physical constraints, and data-driven methods to enable the robust and interpretable discovery of advanced materials.

Description

Keywords

Citation

Source

Digital Discovery

Book Title

Entity type

Publication

Access Statement

License Rights

Restricted until