This technology utilizes an offline reinforcement learning model to derive grasping poses for irregular objects. It extracts the workspace from offline data collected within the robot's operating environment and applies a penalty to actions with high Q-values that are not present in the valid offline dataset, thereby preventing excessive Q-value predictions outside the offline data distribution and optimizing the grasping success rate.
Real-time online reinforcement learning methods carry a high risk of robot damage during data collection, are time-consuming and costly, and are difficult to implement in real-world field applications due to environmental constraints.
This technology employs an offline reinforcement learning-based model to perform training without real-time interaction. By inferring pixel-wise Q-values and applying a penalty to actions with the maximum Q-value that are not included in the existing dataset (valid offline data), the model is prevented from selecting incorrect optimal actions outside the training data distribution, allowing for the precise derivation of 6-DOF grasping poses. It can be applied to logistics picking, service robots, and manufacturing automation, optimizing grasping success rates without the risk of robot damage during data collection.
This invention was developed with the support of the Artificial Intelligence Convergence Innovation Talent Cultivation program by the Ministry of Science and ICT.
N/A