A controller-free robot arm control system where the user selects an object by looking at it through augmented reality glasses, and artificial intelligence identifies the object, calculates its 3D coordinates, and directs the robot arm to approach and retrieve it automatically.

‍

[Device Implementation Example] This image was generated using AI.

‍

Background and Necessity of the Invention

While robot arms are powerful tools that can pick up and move objects in place of human hands, they are typically operated manually via controllers like joysticks or by following pre-set coordinates. The problem is that for those who need this assistance the most—such as patients with quadriplegia who have limited use of their hands—operating a controller is a significant barrier. As a result, their ability to actively control a robot arm has been extremely limited.

To address this, research has been conducted on methods to control robot arms using the user's gaze. In these systems, the user stares at a virtual control screen on a monitor, and sensors track where the user is looking to move the robot arm. However, this method requires the user to constantly shift their focus between the control screen and the robot arm's current position. Furthermore, because the robot arm moves continuously while the gaze is being tracked, there is a risk of collision. Additionally, since the robot arm only operates while the user is staring at the screen, the process is cumbersome and lacks intuitiveness.

What was needed was a method where the user could simply point to an object naturally, and the robot would handle the precise movements on its own—without the need for joysticks, dedicated control screens, or the risk of the robot arm moving erratically.

Two technologies have recently matured to make this possible: augmented reality devices that overlay information onto the real world while sensing depth and gaze, and AI-based object recognition that instantly identifies and labels objects in real-time video. By combining these, a new path is opened where the user selects a target with their gaze, and the AI and augmented reality systems accurately calculate its position and relay it to the robot.

‍

Technical Principles and Implementation Methods

This invention integrates three components: the augmented reality glasses worn by the user, a data processing server that analyzes video, and the robot arm. Metaphorically, the glasses serve as the user's eyes and pointing finger, the server acts as the brain that recognizes objects, and the robot arm functions as the hand that retrieves items. The user simply looks at an object, and the system handles the rest.

First, the glasses transmit the scene the user is viewing to the server as a real-time video feed. The server uses an AI object recognition model to identify objects in the video and sends the results back to the glasses. The glasses then display a bounding box around each recognized object within the user's field of view. This allows the user to see candidate objects, such as cups or bottles, highlighted with boxes right before their eyes.

Selecting an object is simple: the user just needs to look at the box around the desired object for a few seconds. Once the gaze remains on a box for a set duration, the system registers it as a "selection" and changes the color of the box to confirm. There is no need for buttons or joysticks; the target is chosen solely through the user's gaze.

Once an object is selected, the glasses must determine its exact location in space. To do this, the system creates a 3D mesh of the surrounding environment and uses "ray casting"—projecting a virtual beam toward the object—to measure the distance and relative position. This converts the 2D position on the screen into 3D coordinates. By also reading the QR code attached to the robot arm to determine its orientation and current position, the user, the object, and the robot arm are all placed within a single coordinate system.

Once the object's coordinates and the robot arm's position are established, the robot arm automatically approaches the object to retrieve it for the user or return it to its place. The entire process—looking, confirming, calculating coordinates, and retrieving—occurs without the user ever touching a controller. Furthermore, because the system tracks positions in real-time, the margin of error between the calculated coordinates and the actual movement is extremely low.

‍

Advantages and Expected Effects of the Technology

Technology Advantages  The primary strength of this technology is that the robotic arm is controlled solely by the user's gaze. There is no need for joysticks, controllers, or separate control screens to monitor. Users simply look at real-world objects, and the robot handles the precise movement toward the target. Because selection markers are overlaid on the actual scene and the robot calculates its own path, the burden of switching focus between a screen and the robotic arm, as well as the risk of collision, are eliminated. Precise coordinate calculation also ensures the robotic arm reaches its target reliably.

This is fundamentally different from both controller-based and traditional eye-tracking methods. Controllers require a level of manual dexterity that target users may lack, while monitor-based eye tracking often leads to divided attention and collision risks. This technology handles selection through natural gaze in real space, while offloading heavy tasks like object recognition, 3D positioning, and movement execution to AI and robotics. It is built on proven components—such as widely used augmented reality devices, AI recognition models, and QR code-based calibration—ensuring both practicality and reproducibility.

Applications  The most direct application is in daily living assistance and rehabilitation for individuals with limited mobility. Patients with paralysis who cannot use their hands can pick up and move everyday objects using only their gaze, and the system can be mounted on a wheelchair to serve as a personal assistive robotic arm. Beyond caregiving, this approach can be expanded to remote operations that require intuitive, hands-free control.

By enabling users to handle objects independently using only their gaze, individuals with severe physical limitations can regain daily autonomy and significantly improve their quality of life. As augmented reality devices and real-time AI recognition technologies continue to improve and become more affordable, this intuitive, safe, and controller-free control method is expected to become a foundation for service robots across rehabilitation and caregiving sectors.

‍

[Operating Principle] When an augmented reality user selects an object by looking at it, the server (data processing unit) detects the object using AI, converts the coordinates into 3D space, and transmits them to the robotic arm, which then grasps and moves the object.

‍


Patent Listing  IBL-26-0826

Inventors Professor Dong-Joo Kim, Hak-Seung Kim, Se-Ho Lee, Jung-Woo Hyung (Department of Brain and Cognitive Engineering, Korea University)

‍

‍