Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Blind Data Linkage Using n-gram Similarity Comparisons

dc.contributor.authorChurches, Tim
dc.contributor.authorChristen, Peter
dc.coverage.spatialSydney Australia
dc.date.accessioned2015-12-13T22:38:04Z
dc.date.available2015-12-13T22:38:04Z
dc.date.createdMay 26-28 2004
dc.date.issued2004
dc.date.updated2016-02-24T09:48:48Z
dc.description.abstractIntegrating or linking data from different sources is an increasingly important task in the preprocessing stage of many data mining projects. The aim of such linkages is to merge all records relating to the same entity, such as a patient or a customer. If no common unique entity identifiers (keys) are available in all data sources, the linkage needs to be performed using the available identifying attributes, like names and addresses. Data confidentiality often limits or even prohibits successful data linkage, as either no consent can be gained (for example in biomedical studies) or the data holders are not willing to release their data for linkage by other parties. We present methods for confidential data linkage based on hash encoding, public key encryption and n-gram similarity comparison techniques, and show how blind data linkage can be performed.
dc.identifier.isbn0302-9743
dc.identifier.urihttp://hdl.handle.net/1885/77382
dc.publisherSpringer
dc.relation.ispartofseriesPacific Asia Conference on Knowledge Discovery and Data Mining (PAKDD 2004)
dc.sourceAdvances in Knowledge Discovery and Data Mining. 8th Pacific-Asia Conference, PAKDD 2004 Proceedings
dc.source.urihttp://www.springeronline.com/sgw/cda/frontpage/0,11855,1-164-22-30501393-0,00.html?changeHeader=true
dc.subjectKeywords: Customer satisfaction; Data handling; Data mining; Project management; Data linkage; Linking data; Unique entity identifiers; Artificial intelligence Data matching; Hash encoding; N-gram indexing; Privacy preserving data mining; Public key infrastructure
dc.titleBlind Data Linkage Using n-gram Similarity Comparisons
dc.typeConference paper
local.bibliographicCitation.lastpage126
local.bibliographicCitation.startpage121
local.contributor.affiliationChurches, Tim, NSW Health
local.contributor.affiliationChristen, Peter, College of Engineering and Computer Science, ANU
local.contributor.authoruidChristen, Peter, u4021539
local.description.notesImported from ARIES
local.description.refereedYes
local.identifier.absfor080109 - Pattern Recognition and Data Mining
local.identifier.absfor080402 - Data Encryption
local.identifier.ariespublicationMigratedxPub6249
local.identifier.scopusID2-s2.0-7444258692
local.type.statusPublished Version

Downloads