Dong, ZhenKong, YuLiu, CuiweiLi, HongdongJia, Yunde2015-12-10November 2http://hdl.handle.net/1885/64676In this paper, we address the problem of recognizing human interaction of two persons from videos. We fuse global and local features to build a more expressive and discriminative action representation. The representation based on multiple features is robust to motion ambiguity and partial occlusion in interactions. Moreover, action context information is utilized to capture the interdependencies between interaction class and individual action classes of two persons. We introduce a hierarchical random field model which integrates large-scale global feature, local spatial-temporal feature and action context information into a unified framework. Results on UT-Interaction dataset show that our method is quite effective in recognizing human interaction.Keywords: Context information; Data sets; Global feature; Human interactions; Local feature; Multiple features; Partial occlusions; Random field model; Spatial temporals; Unified framework; Pattern recognition; SemanticsRecognizing Human Interaction by Multiple Features201110.1109/ACPR.2011.61665332016-02-24