Automatic identification of the most important elements in an XML collection
| dc.contributor.author | Krumpholz, Alexander | |
| dc.contributor.author | Hadad, Amir | |
| dc.contributor.author | Studeny, Nina | |
| dc.contributor.author | Gedeon, Tamas (Tom) | |
| dc.contributor.author | Hawking, David | |
| dc.coverage.spatial | Canberra Australia | |
| dc.date.accessioned | 2015-12-10T22:20:53Z | |
| dc.date.created | December 2 2011 | |
| dc.date.issued | 2011 | |
| dc.date.updated | 2016-02-24T11:30:38Z | |
| dc.description.abstract | An important problem in XML retrieval is determining the most useful element types to retrieve - e.g. book, chapter, section, paragraph or caption. An automated system for doing this could be based on features of element types related to size, depth, frequency of occurrence, etc. We consider a large number of such features and assess their usefulness in predicting the types of elements judged relevant in INEX evaluations for the IEEE and Wikipedia 2006 corpora. For each feature we automatically assign Useful / Not-Useful labels to element types using Fuzzy c-Means Clustering. We then rank the features by the accuracy with which they predict the manual judgments. We find strong overlap between the top-ten most predictive features for the two collections and that seven features achieve high average accuracy (F-measure > 65%) acrosss them. We hypothesize that an XML retrieval system working on an unlabelled corpus could use these features to decide which retrieval units are most appropriate to return to the user. | |
| dc.identifier.isbn | 9781921426926 | |
| dc.identifier.uri | http://hdl.handle.net/1885/52132 | |
| dc.publisher | RMIT University | |
| dc.relation.ispartofseries | Australasian Document Computing Symposium (ADCS 2011) | |
| dc.source | Proceedings of the Sixteenth Australasian Document Computing Symposium | |
| dc.source.uri | http://www.cs.rmit.edu.au/adcs2011/ | |
| dc.subject | Keywords: Automated systems; Automatic identification; Element type; F-measure; Fuzzy C means clustering; Wikipedia; XML Retrieval; XML retrieval systems; Automation; Fuzzy systems; XML F-Measure; Fuzzy C-Means Clustering; XML Retrieval | |
| dc.title | Automatic identification of the most important elements in an XML collection | |
| dc.type | Conference paper | |
| local.bibliographicCitation.lastpage | 17 | |
| local.bibliographicCitation.startpage | 14 | |
| local.contributor.affiliation | Krumpholz, Alexander, CSIRO | |
| local.contributor.affiliation | Hadad, Amir, College of Engineering and Computer Science, ANU | |
| local.contributor.affiliation | Studeny, Nina, Technical University of Vienna | |
| local.contributor.affiliation | Gedeon, Tamas (Tom), College of Engineering and Computer Science, ANU | |
| local.contributor.affiliation | Hawking, David, College of Engineering and Computer Science, ANU | |
| local.contributor.authoruid | Hadad, Amir, u4366050 | |
| local.contributor.authoruid | Gedeon, Tamas (Tom), u4088783 | |
| local.contributor.authoruid | Hawking, David, a109750 | |
| local.description.embargo | 2037-12-31 | |
| local.description.notes | Imported from ARIES | |
| local.description.refereed | Yes | |
| local.identifier.absfor | 080108 - Neural, Evolutionary and Fuzzy Computation | |
| local.identifier.absfor | 080399 - Computer Software not elsewhere classified | |
| local.identifier.absfor | 100503 - Computer Communications Networks | |
| local.identifier.absseo | 970108 - Expanding Knowledge in the Information and Computing Sciences | |
| local.identifier.ariespublication | u4963866xPUB239 | |
| local.identifier.ariespublication | f5625xPUB7568 | |
| local.identifier.scopusID | 2-s2.0-84872860710 | |
| local.type.status | Published Version |
Downloads
Original bundle
1 - 1 of 1
Loading...
- Name:
- 01_Krumpholz_Automatic_identification_of_2011.pdf
- Size:
- 259.63 KB
- Format:
- Adobe Portable Document Format