Imitation learning is attracting attention as a way to teach robots dexterous tasks such as pinching, twisting and inserting. It collects demonstration data from people performing the task and trains the robot on that data. This article explains how to collect data for imitation learning, focusing on hand data.
What is imitation learning?
Unlike reinforcement learning, which relies on designed rewards and trial and error, imitation learning learns from large numbers of human demonstrations. The more complex the task, the harder it is to design rewards, so showing the robot how to do it is often more efficient. On the other hand, learning quality depends heavily on the quality and quantity of demonstration data.
Why collecting hand manipulation data is hard
- Many degrees of freedom: the human hand has more than 20 degrees of freedom, and every finger joint angle must be recorded accurately
- Touch matters: without knowing where and how hard the hand is touching, a robot cannot reproduce the task
- Real-time response: when collecting data via teleoperation, high latency prevents the operator from moving naturally
Common collection methods
There are three main ways to collect hand manipulation data:
- Camera-based hand pose estimation: easy to set up, but accuracy drops when fingers are occluded and touch cannot be captured
- data glove: sensors worn on the hand directly record finger motion and contact
- Teleoperation: a data glove or similar device drives the robot hand, and the robot's own motion and sensor data are recorded directly
In practice, researchers often capture human hand motion with a data glove and mirror it on a multi-fingered robot hand in real time. Because the robot's own joint and force data become the training data, this approach easily absorbs differences between human and robot bodies.
How to choose a data glove
- Tactile sensor density and coverage: fingertips only or the full palm, number of sensing points and resolution
- Hand position and orientation accuracy: fingertip position accuracy (in mm) and orientation accuracy
- Latency: for teleoperation, aim for roughly 10–20 ms or less from motion to data output
- Software: whether it supports a Python SDK and ROS 2 and fits into your existing training pipeline
- Robot hand integration: whether a mapping to the robot hand you plan to use is available
For example, Wuji Glove has 526 tactile sensing points across the palm and outputs data with fingertip position accuracy of 2 mm or better and latency of 10 ms or less (wired) via electromagnetic tracking. Combined with the 20-DoF Wuji Hand 2 robotic hand, it supports the full loop of operation, reproduction and learning.
- Wuji Glove: tactile data glove (526 tactile sensing points, electromagnetic tracking)
- Wuji Hand 2: 20-DoF robot hand with direct drive on every joint
From data collection to training
- Capture: record human hand motion and touch with the data glove
- Mapping: convert human hand motion to robot hand joints in real time
- Execution: the robot hand reproduces the motion while robot-side data is recorded
- Training: train an imitation learning model on the collected data and validate it in simulation (MuJoCo, Isaac Sim, etc.) and on real hardware
Summary
Imitation learning results depend on the quality of demonstration data. For dexterous manual work in particular, it is crucial to record not only finger motion but also touch. As the authorized distributor of Wuji Technology in Japan, HumansX supports you from selecting data gloves and robot hands to deployment and technical support. Feel free to contact us for a demo.