Sold
Available
IBL-26-0705

Method for learning robot object grasping poses, server for learning robot object grasping poses, and system for learning robot object grasping poses

Listed on
2026-07-24
Robot-related Technology Robot Arm/Manipulator Control/AI/SW
0.07
CI (SI)
★★★★★★★★★★
1.03
TR (N)
★★★★★★★★★★
0.06
MC
★★★★★★★★★★
Grasping Learning System for Irregular Objects Using Offline Reinforcement Learning and Distribution Shift Penalties

This technology utilizes an offline reinforcement learning model to derive grasping poses for irregular objects. It extracts the workspace from offline data collected within the robot's operating environment and applies a penalty to actions with high Q-values that are not present in the valid offline dataset, thereby preventing excessive Q-value predictions outside the offline data distribution and optimizing the grasping success rate.

Real-time online reinforcement learning methods carry a high risk of robot damage during data collection, are time-consuming and costly, and are difficult to implement in real-world field applications due to environmental constraints.

This technology employs an offline reinforcement learning-based model to perform training without real-time interaction. By inferring pixel-wise Q-values and applying a penalty to actions with the maximum Q-value that are not included in the existing dataset (valid offline data), the model is prevented from selecting incorrect optimal actions outside the training data distribution, allowing for the precise derivation of 6-DOF grasping poses. It can be applied to logistics picking, service robots, and manufacturing automation, optimizing grasping success rates without the risk of robot damage during data collection.

Key Features:
  • Offline reinforcement learning model training step, where the training unit of the learning server trains an offline reinforcement learning model using valid offline data and infers pixel-wise Q-values
  • Penalty application step, where the action with the highest Q-value among those not present in the valid offline data is selected as a representative action, and a penalty is applied only to that representative action
  • Maximum Q-value action determination step, where it is verified whether an action is the one with the highest Q-value among those determined to be excluded from the valid offline data
  • Valid offline data extraction step, where valid offline data corresponding to the workspace is extracted

This invention was developed with the support of the Artificial Intelligence Convergence Innovation Talent Cultivation program by the Ministry of Science and ICT.

Hanyang University, ERICA campus
Taejun Park | Seunghwan Yoo | Jongwan Yoon | Byeongjin Ko | Yungi Hong | Juyeol Park
Document
Date of application:
2023-12-28
|
Patent registration number:
10-2939466
Industry
robot•automation
Technology
Robotics
Artifical Intelligence
Country
Korea
Family Patent

N/A

Price
가격협의
Subscribe to our newsletter to receive the latest patent information faster than anyone else.
← Back to list