AI for humanoid robots trained from one million hours of video

Announced by Dyna Robotics (USA) on August 11, DYNA-2 is trained entirely based on videos from a human's personal perspective, instead of robot action data as is being applied by humanoid robot companies. One million hours of footage is equivalent to 171.1 years of continuous waking experience for each person (assuming eight hours of sleep per day is subtracted).
Dyna Robotics' data training method allows robots to learn physical skills, such as how humans interact with objects and their surroundings, through real-life videos. In particular, DYNA-2 is equipped with a world simulation architecture that combines frame prediction and next actions, using human video to "develop understanding", such as how the physical environment changes, how objects react to movement.

This approach is considered to help reduce dependence on manually collected remote control data, thereby providing a more effective way to train robots, even when expanded. In addition to being effective with robots, it also allows robots to "transfer" to other machines.
With the task of high-precision manufacturing, Dyna Robotics said DYNA-2 helps increase the success rate from 20% to 80-90% simply by increasing the new training scale. For example, a humanoid robot only needed 13 minutes of data for its arm to twist open a bottle cap.
The model also demonstrates transferability between fixed robotic arms, humanoid robot prototypes, and robotic hands. According to the developer, on 15 standard tasks, robots running DYNA-2 always gave better results. Dyna Robotics compared the new platform with the previous DYNA-1 model, which used the usual visual-language-action architecture. As a result, DYNA-2 met 87% of the requirements, nearly double the 46% of DYNA-1.
Self-adaptation is also a highlight on DYNA-2. In tests involving daily activities such as chopping food and cleaning the workspace, the DYNA-2 was able to adapt itself to its surroundings, while the old model had to do it manually. With guided tasks, the new video training method helped the robot score 133% higher than the old method.
"For many years, the field of robotics was 'clogged' by the data problem. Manually collecting physical data simply could not scale if it wanted to achieve super intelligence," Jason Ma, co-founder of Dyna Robotics, told Interesting Engineering. "Action data may be scarce, but video is everywhere. With DYNA-2, we demonstrate physical intuition does not require millions of hours of training on a robotic arm, and can instead be learned directly from human video."
Unlike AI that processes text, images or sounds, robots must learn from data associated with movement in real space, including images from cameras, joint states, impact forces, object positions and feedback after each action.
According to TechRadar, this data source is difficult to collect on a large scale because each robot has a different mechanical design, sensors and control system. Letting a robot perform millions of operations on its own is time-consuming and risks damaging the device. Therefore, solutions like Dyna Robotics's are expected to accelerate the training of humanoid robots for real-world environments.
According to ChinaDaily, data is actually a major bottleneck in the field of humanoid robots. In many places, especially in China, many training facilities have been built recently but have not met enough demand, causing the speed of application of this type of robot to the real environment to slow down significantly.
According to China's Ministry of Industry and Information Technology, last year the country had more than 140 manufacturers and more than 330 models of humanoid robots. Meanwhile, statistics from market research company Smart Analytics Global (SAG) show that China accounted for 97% of humanoid robots shipped in the first half of the year, led by Agibot and Unitree.