Visual robotic picking: how to learn from classical control
Abstract
Reaching and grasping are particularly important for robotic manipulation both as standalone tasks and prerequisites for high-order manipulation. A classical engineering approach involves understanding the target object and estimating the grasp/release pose as separate sub-modules from the robot control in the full manipulation pipeline. An appealing alternative is to compute the control from raw image pixels directly. In the established image-based visual servo control literature, this paradigm is known to lead to more robust and responsive systems in real-world scenarios. However, existing visual servoing algorithms cannot deal with complex, unstructured, and dynamic real-world environments with multiple instances of target objects, high levels of clutter, or requiring path planning and tracking logic as part of the high-level control algorithm. This thesis makes several advances to the field of visuomotor (image-to-control) learning for picking in unstructured, dynamic and multi-instance environments and demonstrates the application of these ideas to human-robot collaboration tasks. I propose and develop a general architecture, the Lyapunov Reaching Network (LyRN), for the design of visual reaching and grasping. LyRN addresses the challenge of combining the decision logic associated with choosing between different target instances and the low-level control of visuomotor reaching and grasping into a single integrated closed-loop control algorithm. LyRN is based on a deep neural network architecture that learns a mapping from the image space to a formulated compound control Lyapunov function (cLf). This compound cLf can be conceptualised as an energy map that encodes the complete set of actions a robot may take for a given scene and robot configuration. The reactive multiinstance capability emerges naturally by following the action associated with the current minimum energy. This holistic view of the environment and the robot motion results in naturally highly robust and reactive systems. A future intelligent robot also needs to be deployed in shared environments, understand human intents and collaborate with humans. In the final part of the thesis, I develop and demonstrate a novel integrated vision-based human-robot collaborative system in an assembly scenario. The robot task is to select a suitable target object from a multi-instance cluttered collection based on an understanding of the task the human is undertaking. The robot must then robustly pick up the object and pass it to the human at the appropriate time for their task. The reaching and grasping algorithm is based on the LyRN architecture, while the semantic understanding is drawn from a parallel work at the Australian National University. The integrated system can deal with unstructured and dynamic environments, perceive and understand human action, and respond based on human action and task progress. This work further demonstrates the reactiveness and robustness of LyRN and shows the potential of leveraging implicit semantic communication for efficient and seamless human-robot collaboration.
Description
Keywords
Citation
Collections
Source
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
Downloads
File
Description
Thesis Material