Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Transforming pairwise duplicates to entity clusters for high-quality duplicate detection

dc.contributor.authorDraisbach, Uwe
dc.contributor.authorChristen, Peter
dc.contributor.authorNaumann, Felix
dc.date.accessioned2023-07-20T00:08:00Z
dc.date.issued2019
dc.date.updated2022-05-22T08:15:48Z
dc.description.abstractDuplicate detection algorithms produce clusters of database records, each cluster representing a single real-world entity. As most of these algorithms use pairwise comparisons, the resulting (transitive) clusters can be inconsistent: Not all records within a cluster are sufficiently similar to be classified as duplicate. Thus, one of many subsequent clustering algorithms can further improve the result. We explain in detail, compare, and evaluate many of these algorithms and introduce three new clustering algorithms in the specific context of duplicate detection. Two of our three new algorithms use the structure of the input graph to create consistent clusters. Our third algorithm, and many other clustering algorithms, focus on the edge weights, instead. For evaluation, in contrast to related work, we experiment on true real-world datasets, and in addition examine in great detail various pair-selection strategies used in practice. While no overall winner emerges, we are able to identify best approaches for different situations. In scenarios with larger clusters, our proposed algorithm, Extended Maximum Clique Clustering (EMCC), and Markov Clustering show the best results. EMCC especially outperforms Markov Clustering regarding the precision of the results and additionally has the advantage that it can also be used in scenarios where edge weights are not available.en_AU
dc.format.mimetypeapplication/pdfen_AU
dc.identifier.issn1936-1955en_AU
dc.identifier.urihttp://hdl.handle.net/1885/294441
dc.language.isoen_AUen_AU
dc.publisherAssociation for Computing Machinary, Inc.en_AU
dc.rights© 2019 Association for Computing Machinery.en_AU
dc.sourceJournal of Data and Information Qualityen_AU
dc.titleTransforming pairwise duplicates to entity clusters for high-quality duplicate detectionen_AU
dc.typeJournal articleen_AU
local.bibliographicCitation.issue1en_AU
local.bibliographicCitation.lastpage30en_AU
local.bibliographicCitation.startpage1en_AU
local.contributor.affiliationDraisbach, Uwe, University of Potsdamen_AU
local.contributor.affiliationChristen, Peter, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationNaumann, Felix, University of Potsdamen_AU
local.contributor.authoruidChristen, Peter, u4021539en_AU
local.description.embargo2099-12-31
local.description.notesImported from ARIESen_AU
local.identifier.absfor460504 - Data qualityen_AU
local.identifier.absfor460502 - Data mining and knowledge discoveryen_AU
local.identifier.absfor460507 - Information extraction and fusionen_AU
local.identifier.ariespublicationa383154xPUB11539en_AU
local.identifier.citationvolume12en_AU
local.identifier.doi10.1145/3352591en_AU
local.identifier.scopusID2-s2.0-85077792876
local.identifier.thomsonIDWOS:000535150100003
local.publisher.urlhttps://dl.acm.org/en_AU
local.type.statusPublished Versionen_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Transforming Pairwise Duplicates to Entity Clusters.pdf
Size:
1.08 MB
Format:
Adobe Portable Document Format
Description: