Spatial encoding of visual words for image classification
| dc.contributor.author | Liu, Dong | |
| dc.contributor.author | Wang, Shengsheng | |
| dc.contributor.author | Porikli, Fatih | |
| dc.date.accessioned | 2016-09-02T03:50:33Z | |
| dc.date.available | 2016-09-02T03:50:33Z | |
| dc.date.issued | 2016-05-31 | |
| dc.description.abstract | Appearance-based bag-of-visual words (BoVW) models are employed to represent the frequency of a vocabulary of local features in an image. Due to their versatility, they are widely popular, although they ignore the underlying spatial context and relationships among the features. Here, we present a unified representation that enhances BoVWs with explicit local and global structure models. Three aspects of our method should be noted in comparison to the previous approaches. First, we use a local structure feature that encodes the spatial attributes between a pair of points in a discriminative fashion using class-label information. We introduce a bag-of-structural words (BoSW) model for the given image set and describe each image with this model on its coarsely sampled relevant keypoints. We then combine the codebook histograms of BoVW and BoSW to train a classifier. Rigorous experimental evaluations on four benchmark data sets demonstrate that the unified representation outperforms the conventional models and compares favorably to more sophisticated scene classification techniques. | en_AU |
| dc.description.sponsorship | This work was supported under the Australian Research Council’s Discovery Projects funding scheme (Project No. DP150104645) and the National Natural Science Foundation of China (No. 61472161). | en_AU |
| dc.identifier.issn | 1017-9909 | en_AU |
| dc.identifier.uri | http://hdl.handle.net/1885/108600 | |
| dc.publisher | Society of Photo-optical Instrumentation Engineers (SPIE) | en_AU |
| dc.relation | http://purl.org/au-research/grants/arc/DP150104645 | en_AU |
| dc.rights | http://www.sherpa.ac.uk/romeo/issn/1017-9909/..."Publisher's version/PDF may be used (preferred)" from SHERPA/RoMEO site (as at 2/09/16). | en_AU |
| dc.rights | Copyright 2016 2016 SPIE and IS&T. One print or electronic copy may be made for personal use only. Systematic reproduction and distribution, duplication of any material in this paper for a fee or for commercial purposes, or modification of the content of the paper are prohibited. The full citation of the paper: Liu, Dong, Shengsheng Wang, and Fatih Porikli. "Spatial encoding of visual words for image classification." Journal of Electronic Imaging 25.3 (2016): 033008-033008. | en_AU |
| dc.source | Journal of Electronic Imaging | en_AU |
| dc.subject | visual descriptors | en_AU |
| dc.subject | bag-of-words | en_AU |
| dc.subject | spatial feature representations | en_AU |
| dc.subject | scene classification | en_AU |
| dc.title | Spatial encoding of visual words for image classification | en_AU |
| dc.type | Journal article | en_AU |
| dcterms.accessRights | Open Access | en_AU |
| local.bibliographicCitation.issue | 3 | en_AU |
| local.bibliographicCitation.startpage | 033008 | en_AU |
| local.contributor.affiliation | Porikli, F., Research School of Engineering, The Australian National University | en_AU |
| local.contributor.authoruid | u5405232 | en_AU |
| local.identifier.citationvolume | 25 | en_AU |
| local.identifier.doi | 10.1117/1.JEI.25.3.033008 | en_AU |
| local.publisher.url | http://spie.org/ | en_AU |
| local.type.status | Published Version | en_AU |