Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Active Learning Based Similarity Filtering for Efficient and Effective Record Linkage

Loading...
Thumbnail Image

Date

Authors

Nanayakkara, Charini
Christen, Peter
Ranbaduge, Thilina

Journal Title

Journal ISSN

Volume Title

Publisher

Springer Nature Switzerland AG

Abstract

The limited analytical value of using individual databases on their own increasingly requires the integration of large and complex databases for advanced data analytics. Linking personal medical records with travel and immigration data, for example, will allow the effective management of pandemics such as the current COVID-19 outbreak by tracking potentially infected individuals and their contacts. One major challenge for accurate linkage of large databases is the quadratic or even higher computational complexities of many advanced linkage algorithms. In this paper we present a novel approach that, based on the expected number of true matches between two databases, applies active learning to remove compared record pairs that are likely non-matches before a computationally expensive classification or clustering algorithm is employed to classify record pairs. Unlike blocking and indexing techniques that are used to reduce the number of record pairs to be compared, using recursive binning on a data dimension such as time or space, our approach removes likely non-matching record pairs in each bin after their comparison. Experiments on two real-world databases show that similarity filtering can substantially reduce run time and improve precision, at the costs of a small reduction in recall, of the final linkage results.

Description

Citation

Source

Lecture Notes in Artificial Intelligence

Book Title

Advances in Knowledge Discovery and Data Mining : 25th Pacific-Asia Conference, PAKDD 2021 Virtual Event, May 11–14, 2021 Proceedings, Part II

Entity type

Access Statement

Open Access

License Rights

Restricted until

Downloads