Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Robust temporal graph clustering for group record linkage

dc.contributor.authorNanayakkara, Charini
dc.contributor.authorChristen, Peter
dc.contributor.authorRanbaduge, Thilina
dc.contributor.editorGong, Z
dc.contributor.editorHuang, S-J
dc.contributor.editorZhou, Z-H
dc.contributor.editorZhang, M-L
dc.contributor.editorYang, Q
dc.coverage.spatialMacua, China
dc.date.accessioned2024-02-12T00:41:24Z
dc.date.createdApril 14-17 2019
dc.date.issued2019
dc.date.updated2022-10-02T07:19:24Z
dc.description.abstractResearch in the social sciences is increasingly based on large and complex data collections, where individual data sets from different domains need to be linked to allow advanced analytics. A popular type of data used in such a context are historical registries containing birth, death, and marriage certificates. Individually, such data sets however limit the types of studies that can be conducted. Specifically, it is impossible to track individuals, families, or households over time. Once such data sets are linked and family trees are available it is possible to, for example, investigate how education, health, mobility, and employment influence the lives of people over two or even more generations. The linkage of historical records is challenging because of data quality issues and because often there are no ground truth data available. Unsupervised techniques need to be employed, which generally are based on similarity graphs generated by comparing individual records. In this paper we present a novel temporal clustering approach aimed at linking records of the same group (such as all births by the same mother) where temporal constraints (such as intervals between births) need to be enforced. We combine a connected component approach with an iterative merging step which considers temporal constraints to obtain accurate clustering results. Experiments on a real Scottish data set show the superiority of our approach over a previous clustering approach for record linkage.en_AU
dc.description.sponsorshipThis work was supported by ESRC grants ES/K00574X/2 Digitising Scotland and ES/L007487/1 ADRC-S. We like to thank Alice Reid of the University of Cambridge and her colleagues Ros Davies and Eilidh Garrett for their work on the Isle of Skye database, and their helpful advice on historical Scottish demography. This work was partially funded by the Australian Research Council under DP160101934.en_AU
dc.format.mimetypeapplication/pdfen_AU
dc.identifier.isbn978-303016144-6en_AU
dc.identifier.urihttp://hdl.handle.net/1885/313379
dc.language.isoen_AUen_AU
dc.provenancehttps://www.springer.com/gp/rights-permissions/our-policy-on-archiving-in-institutional-or-funding-body-reposit/6629030..."The Accepted Version can be archived in a Non-Commercial Institutional Repository. 24 months embargo" from SHERPA/RoMEO site (as at 14/2/2024).
dc.publisherSpringeren_AU
dc.relationhttp://purl.org/au-research/grants/arc/DP160101934en_AU
dc.relation.ispartofseries23rd Pacific-Asia Conference on Knowledge Discovery and Data Mining, PAKDD 2019en_AU
dc.rights© 2019 Springer Nature Switzerland AG 2019en_AU
dc.sourceLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)en_AU
dc.subjectEntity resolutionen_AU
dc.subjectStar clusteringen_AU
dc.subjectVital recordsen_AU
dc.subjectBirth bundlingen_AU
dc.titleRobust temporal graph clustering for group record linkageen_AU
dc.typeConference paperen_AU
dcterms.accessRightsOpen Access
local.bibliographicCitation.lastpage538en_AU
local.bibliographicCitation.startpage526en_AU
local.contributor.affiliationNanayakkara, Charini, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationChristen, Peter, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationRanbaduge, Thilina, College of Engineering and Computer Science, ANUen_AU
local.contributor.authoruidNanayakkara, Charini, u6507558en_AU
local.contributor.authoruidChristen, Peter, u4021539en_AU
local.contributor.authoruidRanbaduge, Thilina, u5421298en_AU
local.description.notesImported from ARIESen_AU
local.description.refereedYes
local.identifier.absfor460507 - Information extraction and fusionen_AU
local.identifier.absfor460504 - Data qualityen_AU
local.identifier.absfor460502 - Data mining and knowledge discoveryen_AU
local.identifier.ariespublicationu3102795xPUB1601en_AU
local.identifier.doi10.1007/978-3-030-16145-3_41en_AU
local.identifier.scopusID2-s2.0-85064947538
local.identifier.thomsonIDWOS:000716968700041
local.publisher.urlhttps://link.springer.com/en_AU
local.type.statusAccepted Versionen_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
nanayakkara2019clustering.pdf
Size:
432.1 KB
Format:
Adobe Portable Document Format